Imported from amirulhazym/hermes-agent-personal_assistant (
skills/med-tracker/SKILL.md). Install upstream withnpx skills add amirulhazym/hermes-agent-personal_assistant --skill med-tracker. Copyright stays with the author.
Medication Tracker
⚠️ MANDATORY: Resolve Before Confirming (DO NOT SKIP)
This section MUST be followed. Violation caused a critical fabrication incident on 2026-07-04.
When you see ANY drug name or shorthand in a user message:
- DO NOT assume you know what it means — even if it sounds familiar
- DO NOT fabricate a mapping based on the name sounding similar to another drug
- FIRST check this skill's Drug-Level Confirmation Patterns table below
- SECOND run
python3 med_resolve.py <drug_name> [--time HH:MM] [--slot LETTER] - THIRD only use the
drug_idreturned by the resolve script - If resolve returns
UNKNOWN, ask the user which drug they mean
Quoted Transcript Guard
Inbound messages containing pasted WhatsApp transcript lines such as [24/07, 6:00 am] Sender: are discussion/history, not fresh medication confirmations. The gateway hook must reject the entire message before scanning completion words or drug names inside quoted history. Regression test: test_quoted_whatsapp_transcript_is_not_confirmation.
Scenario from 2026-07-04 (DO NOT REPEAT):
- User said "letram" (standard shorthand for Levetiracetam)
- Agent fabricated: "Letrozole" (zero basis, no source, pure hallucination)
- Root cause: agent didn't load this skill, didn't check the disambiguation table, didn't search past sessions for "letram = Levetiracetam"
- Result: wrong drug logged, user furious, trust damaged
Always remember: The disambiguation table below maps "letram" → levetiracetam_b/levetiracetam_e.
Never fabricate a mapping if it's not in this table.
Critical Truth: Medication Priority Hierarchy
This is the single most important rule in the entire med system. Every timing decision, gap calculation, reminder logic, and chain display derives from this hierarchy. Violating it causes the user to explode.
Priority order (highest → lowest):
- Dexamethasone — the taper drives the entire schedule. Gaps B→C and C→D exist SOLELY for Dexa spacing. Supplements MUST NOT shift this timing.
- Akurit-4 — anchors the morning chain. Empty-stomach requirement is non-negotiable.
- Levetiracetam — must maintain ~12h gap between morning (B) and evening (E) doses. Takes timing from B.
- Calcium + Calcitriol — passengers on Slot C. Their intake time MUST NOT affect downstream Dexa timing (D).
- B-Complex — optional (Rabu/Sabtu only). Never influences chain timing.
Practical consequence: In get_actual_time(), if a slot contains Dexamethasone, return that drug's time — NOT the latest time among all drugs in the slot. Taking Calcium at 13:00 after Dexa at 12:15 MUST NOT push D ready to 17:00. D stays at 16:15.
User's words (verbatim, 2026-07-07): "OUR MAIN PRIORITY IS DEXA, AKURIT-4 DAN LETRAM. ANY GAPS OR TIMEFRAME MUST PRIORITIZE THIS, BUKAN CC."
Architecture (v3 — 2026-07-05)
Layer 1: cron (chain_monitor.sh) → Every 15 min, state-aware reminders (no_agent=true)
Layer 1b: cron (taper_alert.py) → Daily 06:00, tapering phase transition alerts
Layer 2: chat agent (THIS SKILL) → Detect confirmations, log, show chain
Layer 3: engine (chain_calc.py) → Read schedule + status + taper, calculate chain
All layers share these data files:
med-schedule.json— drug rules, gaps, windows (static)med-status.json— drug-level intake log (written by user confirmation)chain-state.json— reminder counts + cooldown timestamps +todaydate for day-boundary resetdexa_taper.json— dexamethasone tapering schedule (date-dependent dosing)med-supply.json— pill inventory per drug (untracked by default:current: null; NO auto-decrement on confirmation)substitutions.json— drug substitution databasemed-interactions.json— drug interaction safety data
Slot Override System (v3.1 — 2026-07-08)
When user explicitly says "stick with original schedule" / "guna jadual asal" after an early/late outlier dose, the chain-calculated shift must be suppressed.
Mechanism: slot_overrides in chain-state.json:
{
"slot_overrides": {
"2026-07-08": {
"B": {
"suppress_until": "08:00",
"reason": "User said stick with original schedule after early A at 04:04"
}
}
}
}
How to set (when user says "stick with original schedule"):
python3 -c "
import json, datetime
from pathlib import Path
state_file = Path.home() / '.hermes' / 'chain-state.json'
state = json.loads(state_file.read_text())
date = datetime.date.today().isoformat()
state.setdefault('slot_overrides', {})
state['slot_overrides'][date] = {
'B': {'suppress_until': '08:00', 'reason': 'User said stick with original schedule'}
}
state_file.write_text(json.dumps(state, indent=2))
"
When to set: Whenever user says any variation of "stick with original/jadual asal" or "ikut timing biasa" after an outlier dose that would shift the chain. Common trigger: early morning A (4-5am) when user plans to sleep and resume normal schedule.
Auto-clear: Day-boundary reset in chain_monitor.sh removes slot_overrides entirely on new day. Also, at suppress_until time, the slot becomes eligible again.
Related code in chain_calc.py (line ~572-590):
today_str = datetime.now(MYT).strftime('%Y-%m-%d')
slot_overrides = chain_state.get('slot_overrides', {}).get(today_str, {})
# ...inside the slot loop:
if slot in slot_overrides:
override = slot_overrides[slot]
suppress_until = override.get('suppress_until')
if suppress_until and now_min < time_str_to_minutes(suppress_until):
continue
Pitfall — Cron isolation: chain_monitor.sh has no conversation context; chat instructions such as sleep windows, original timing, or changed intended time must be translated immediately into chain-state.json overrides. The cron will not infer them. Update the override when the user wakes or changes the intended time, and suppress reminders during an explicit sleep window. | Rationale |
|-------|----------|-----------|
| 0 (first) | 0 min | Fire immediately when slot becomes ready |
| 1-2 | 60 min | Gentle spacing — user gets ~1/hour |
| 3-6 | 30 min | Escalating, but still respectful |
| 7+ | 15 min | Critical — user might have forgotten entirely |
Pitfall: After a gateway restart or cron pause/resume, the last_reminder_times persist in chain-state.json so cooldown survives restarts. Reset the JSON file only when explicitly clearing state.
Pitfall: Don't Manually Test Against Live State
NEVER run chain_monitor.sh (the cron delivery script) against the real chain-state.json. The script writes reminder counts and timestamps as a side effect — running it manually poisons the state. The next real cron tick will see false counts and skip legitimate reminders.
Safe test method:
# Use --next only — reads state without writing
python3 chain_calc.py --next
# To preview output text without recording:
python3 chain_calc.py --next | grep should_fire
python3 chain_calc.py --template --slot D
# Never:
# bash chain_monitor.sh ← WRITES TO LIVE STATE
Why it matters (2026-07-04): Manual bash chain_monitor.sh at 17:31 set D: count=1, last=17:31 in the real state file. The 17:45 cron tick saw cooldown active (only 14 min since last) and went [SILENT]. It looked like cron was broken — it wasn't. The state was polluted by my own test. Wasted 45 minutes debugging a problem I created.
Code entry points: is_within_cooldown() + get_cooldown_interval() in chain_calc.py. The function is called in all three fire paths (partial, regular, default-fallback).
Verification recipe:
from chain_calc import is_within_cooldown, get_cooldown_interval
chain_state = {'last_reminder_times': {'D': '16:45'}}
reminder_counts = {'D': 1}
# 15 min after last → should return True (skip)
is_within_cooldown('D', reminder_counts, chain_state, 17*60+0)
# 60 min after last → should return False (allow fire)
is_within_cooldown('D', reminder_counts, chain_state, 17*60+45)
CRITICAL PITFALL: Day-Boundary Reset for Reminder Counts (v2.1 — 2026-07-07)
Reminder counts in chain-state.json are DAY-SCOPED, not perpetual. Without a day-boundary reset, yesterday's count of 3 for slot E bleeds into today, making the first reminder at 20:00 fire with count=4 template ("dah 4x tanya") when the user hasn't been asked ONCE today. This caused user fury on 2026-07-07.
Mechanism in chain_monitor.sh Step 3:
# Day boundary: reset ALL counts if date changed
today = datetime.date.today().isoformat()
state_today = state.get('today')
if state_today != today:
state['reminder_counts'] = {}
state['last_reminder_sent'] = {}
state['last_reminder_times'] = {}
state['today'] = today
This runs BEFORE the slot-confirmation reset and BEFORE the increment, so counts are always fresh for the current day. The today field is written to chain-state.json and persists across gateway restarts.
Pitfall — stale today field: If chain-state.json was written before the day-boundary feature existed, the file won't have a today field. On first run, state_today will be None, which != today, so the reset triggers once (clearing stale counts) and writes today. After that, the boundary is self-maintaining.
Pitfall — cron runs during midnight cross-over: If a cron tick fires at 23:55 and the next fires at 00:05 (next day), the 00:05 tick will detect state_today != today, reset all counts, and start fresh. No slot should fire reminders across midnight anyway (monitor stops at 22:00, resumes at 05:00), so this edge case is theoretical.
Verification:
# Before fix: yesterday's E=3 bleeds into today
cat chain-state.json | python3 -c "import sys,json; d=json.load(sys.stdin); print(d.get('reminder_counts',{}).get('E','MISSING'))"
# Should show nothing (empty) or 0 if day just rolled over
Design Principle: Hybrid Intelligence
DO NOT treat the cron timer as a 2010-era bot that fires fixed reminders at wall-clock times. The system has LLM capability — use it.
The current Layer 1 (chain_monitor.sh + chain_calc.py) works well for timing math (gap calculations, chain shifts) — pure Python is fast, free, and deterministic. But the decision logic (should a reminder fire right now?) and tone/intelligence of the reminder text should involve the LLM.
Target architecture — Approach C (Hybrid):
- Python (chain_calc.py) calculates ready times via gap rules — reliable, zero-cost
- When Python says "ready to fire" → also checks cooldown (is_within_cooldown) so we don't spam the same slot every 15 min
- When Python says "ready to fire AND cooldown expired" → optionally call LLM API for contextual reminder text that accounts for shift severity, user's current state, tone appropriate to the situation
- When Python says "too early / waiting / within cooldown" → silent (no message). The chain calc is trusted for this decision.
Pitfall — LLM bypass: Never use no_agent=true + script + wall-clock defaults for the REMINDER DECISION. The recurring bug was chain_calc.py falling back to default times (12:00, 16:00, 20:00) when the chain-calculated ready_time said "too early." See references/chain-calc-bug.md for the full trace. This caused 9+ spam reminders in a single day for a slot that wasn't due for another 43 minutes.
Checking before acting: When a user reports wrong timing ("reminder fired too early"), do NOT propose a bandaid fix to the script. Trace the full decision chain: cron → script → chain_calc.py → decide → output. The root cause is almost always in chain_calc.py's fallback logic using wall-clock defaults instead of chain-calculated times.
Core Change (v2 — Drug-Level Tracking)
OLD behaviour: Slot B ✅ = both Dexa AND Levetiracetam assumed taken. Reminders stop.
NEW behaviour: Each drug is tracked individually. Slot B shows:
- ◐ partial if only 1/2 drugs taken → reminders STILL fire for pending drugs
- ✅ completed only when ALL required drugs in that slot are taken
- ⏳ pending if none taken
Status icons in chain display:
✅ 07:00— all drugs in slot taken◐ 08:16— some drugs taken, some still pending~12:16— estimated ready time, nothing taken yet
Drug-Level Confirmation Patterns
When user specifies a drug name instead of slot letter, map to the correct slot + drug_id:
| User says | Maps to | Slot |
|---|---|---|
| "akurit", "akurit-4", "rifampicin" | akurit_4 |
A |
| "pyridoxine", "vitamin B6", "b6" | pyridoxine |
A |
| "letram", "levetiracetam", "levetiracetam pagi" | levetiracetam_b |
B (before 2pm) |
| "dexa", "dexamethasone", "dexamethasone pagi", "steroid pagi" | dexamethasone_1 |
B (before 11am) |
| "dexa", "dexamethasone tengahari", "steroid tengahari" | dexamethasone_2 |
C (11am-2pm) |
| "dexa", "dexamethasone petang", "steroid petang" | dexamethasone_3 |
D (after 4pm) |
| "calcium", "kalsium" | calcium |
C |
| "calcitriol", "vitamin D" | calcitriol |
C |
| "b-complex", "swisse", "vitamin B" | b_complex |
C (Rabu/Sabtu only) |
| "letram mlm", "levetiracetam mlm", "levetiracetam malam" | levetiracetam_e |
E (after 7pm) |
Time-based disambiguation for drugs in multiple slots:
- If user says just "dexa" (no time context), check current MYT time:
- Before 10:30 →
dexamethasone_1(B) - 10:30-14:30 →
dexamethasone_2(C) - After 16:00 →
dexamethasone_3(D)
- Before 10:30 →
- If user says just "letram" / "levetiracetam" (no time context):
- Before 14:00 →
levetiracetam_b(B) - After 14:00 →
levetiracetam_e(E)
- Before 14:00 →
How to Log — Drug-Level
# Mark ALL drugs in a slot as taken (user says "dah makan B")
python3 ~/.hermes/scripts/med_confirm.py B
# Mark a SINGLE drug in a slot (user says "dah makan dexa pagi")
python3 ~/.hermes/scripts/med_confirm.py B dexamethasone_1
# Fuzzy drug name match (user says "dexa")
python3 ~/.hermes/scripts/med_confirm.py B dexa
# With specific time
python3 ~/.hermes/scripts/med_confirm.py B dexamethasone_1 --at 08:16
# CRITICAL: when using slot-level --at, ALWAYS pair with --source-text so the
# verification gate runs. Without it, confirm_slot() skips the gate and marks
# ALL drugs in the slot taken with zero check (see "CRITICAL PITFALL:
# --at <slot> Mode Skips Verification Gate" above).
python3 ~/.hermes/scripts/med_confirm.py B --at 08:16 --source-text "dah makan B jam 8.16"
# Query
python3 ~/.hermes/scripts/med_confirm.py --status # all slots today
python3 ~/.hermes/scripts/med_confirm.py --check B # slot B status
python3 ~/.hermes/scripts/med_confirm.py --check B dexamethasone_1 # single drug
Detection Patterns
Match these patterns (case-insensitive) in user messages:
Slot-level (mark all drugs in slot):
dah makan [A-E]→ log with current timedah makan [A-E] pukul HH:MM→ log with specified timedah makan ubat [A-E]sudah makan [A-E][A-E] done,[A-E] siapate [A-E],took [A-E]/confirm [A-E]dah selesaikan [A-E]/dah selesai [A-E]/[A-E] selesaidah selesaikan [drug_name]/dah selesai [drug_name]→ resolve to slot+timedexa dose petang dah selesaikan/dexa petang done→ D with time context- Any message containing BOTH a drug/slot reference AND a past-tense completion word ("done", "selesai", "selesaikan", "dah ambil", "dah telan") AND a time reference → MUST run med_confirm.py immediately
- Bare
dah makan [time]/thanks remindwith NO drug/slot word → ASK first (food vs ubat). Do not assume food. Do not silent-log. Seereferences/makan-ambiguity-ask-dont-assume.md.
Drug-level (mark single drug):
dah makan [drug_name]→ resolve drug_id + slot via table abovedah [drug_name]→ same resolutionmakan [drug_name]→ same resolutiondah makan [drug_name] pagi/tengahari/petang/mlm→ time-context-awaredah makan [drug_name] pukul HH:MM→ with specific time
Multiple in one message:
dah makan A dan B→ confirm both, separate log callsdah makan dexa pagi and letram→ confirm dexamethasone_1 + levetiracetam_bA and dexa done→ confirm A (slot-level) + dexamethasone_1 (drug-level)
Skip patterns:
skip [A-E]→ do NOT log as confirmedtak makan [A-E]→ same as skipskip [drug_name]→ skip single drug in slot
Wake-up patterns (morning context):
baru bangun,just woke up,dah bangun→ user just woke up- Med A is the relevant slot (taken after waking for solat subuh)
- Don't ask about other slots yet — they're not due
- Acknowledge naturally, mention it's time for A if applicable
Procedure After Detection
- Extract the slot letter, drug name, and optional time from user message
- Log via
med_confirm.pywith appropriate args (drug-level or slot-level) - Show chain — run
python3 ~/.hermes/scripts/chain_calc.py --displayand output the result (chat responses only — NOT in cron reminders). PITFALL:chain_calc.pylives in~/.hermes/scripts/ROOT — it is NOT inmed_chain/(that dir has chain_trace/chain_review/chain_consistency/validate_semantic but no chain_calc). Runningpython3 med_chain/chain_calc.pyfails with Errno 2 (no such file). Same full-path rule applies to step 4's--update <LETTER>. - Reset reminder count for that slot — run
chain_calc.py --update <LETTER>so cron moves to next pending slot
Critical: When to reset (2026-07-07): Whenever user confirms a drug in a slot that goes from pending → partial (some drugs taken, others still pending), ALWAYS reset the reminder count. This bypasses the cooldown and lets the next cron tick fire a NEW reminder about the remaining drugs. Without the reset, the cooldown (60 min after the first reminder) blocks ANY follow-up — the user gets no prompt about remaining drugs because the system is still waiting out the cooldown from the initial reminder.
DEXA TAPER BD UNDERDOSING DEFECT (P0 CLINICAL — VERIFIED 2026-07-07): In chain_calc.py, get_dexa_dose_for_slot() maps Slot C → dose_midday unconditionally (line ~222). During BD taper phases (Phase 10-16, starts 2026-09-09), dose_midday=0 AND Slot D (afternoon) is deactivated in active_slots_by_freq['BD'] = ['A','B','C','E']. Result: 4mg afternoon dose is silently dropped → system calculates 6mg/day vs prescribed 10mg = 4mg/day deficit for ~3 months (Sept-Dec 2026). FIX: when freq=='BD', Slot C must map to dose_2pm (4mg), not dose_midday. Verify by mocking cc.today_myt = lambda: '2026-09-15' then get_dexa_dose_for_slot('C') must return 4 and B+C+E total = 10. NOT YET FIXED — still open, must fix before 2026-09-09.
Do NOT reset when: Slot is already fully completed (overall=completed). The next slot's reminder system handles itself.
Reminder Output Format (Cron Delivery)
The cron script (chain_monitor.sh) uses --template from chain_calc.py to generate text, but the final delivery envelope is controlled by chain_monitor.sh and follows strict formatting rules:
Format Rules
- Natural human-to-human text only. No structured formatting, no headers, no bullet lists, no numbered steps. The template text reads like a chat message from Jane.
- No "---" separator. The "---" divider and anything after it (chain display, reply instructions, medication footnotes) is STRICTLY FORBIDDEN in delivered output.
- No reply instructions. Never tell the user "Reply 'dah makan X'" or "Reply to confirm" in the reminder text. The user already knows the convention.
- No progress fractions for partial slots. Do not say "Baru 3/3 je" or "2/3 required" — this number may be misleading when optional drugs (B-Complex) are counted alongside required ones. Say instead: "Dexa #2 je yang tinggal."
- Last line = log code only. The final line of every delivered reminder is the log code — nothing else after it.
Log Code System
Format: [<SLOT>:<N>-<YYMMDD>]
[C:1-260703] → First reminder for C on 2026-07-03
[B:3-260703] → Third reminder for B on 2026-07-03
SLOT= letter A-EN= 1-indexed reminder count for that slot that day (after increment)YYMMDD= date in 2-digit format (26=2026, 07=July, 03=day)
Purpose: Acts as a cross-reference fingerprint. The backend stores the raw reminder counts in chain-state.json indexed by slot. If the user ever asks "reminder [C:3-260703] ada apa?" or if something seems off with a particular day's reminders, the log code lets Jane look up the exact reminder instance in backend logs to reconstruct context.
Example Output (cron fires for C)
Boss, dah makan Dexa dose tengah hari ke belum? B tadi 08:16 ✅. Nak confirmkan je sebab kau belum reply. Dah pukul 12:39.
[C:1-260703]
Chat Response vs Cron Reminder — Distinction
| Aspect | Chat response (user confirms) | Cron reminder (automated) |
|---|---|---|
| Chain display | ✅ Include: A ✅ 07:00 → B ✅ 08:16 → C ~12:16 |
❌ Not included |
| Reply instructions | ❌ Not included | ❌ Not included |
| Log code | ❌ Not needed | ✅ Last line |
| Format | Brief status + chain | Natural text + log code |
Response Style
- Keep tool calls hidden; acknowledge directly in concise Malay/Manglish.
- After confirmation, show status + chain immediately; do not ask the user to repeat information already given.
- If “makan” is ambiguous, ask once whether it means medicine or food; never assume.
- For CLI result interpretation and future-intent splitting, see
references/20260729-cli-confirmation-and-future-intent.md.
Confirmation CLI and Future Intent
A valid JSON payload can accompany exit code 1 when a slot is partial (overall: partial, confirmed: false). Inspect the payload before classifying the operation as failed, and do not chain partial checks with &&. Log only completed intake; future intent such as “jap lagi makan letram” is not a confirmation. If a natural name is UNKNOWN but the resolver provides a canonical suggestion, resolve that canonical ID before confirming.
- Full completion:
✅ B logged (Dexa + Levetiracetam at 08:30). Chain: A ✅ 07:00 → B ✅ 08:16 → C ~12:16 → D ~16:00 → E ~20:16 - Partial completion:
◐ B partial — Dexa ✅ at 08:16, Levetiracetam still pending. Chain: A ✅ 07:00 → B ◐ 08:16 → C ~12:16 → D ~16:00 → E ~20:16 - No celebration — log + chain, that's it
- If drug name is ambiguous (e.g. "dah makan dexa" at 1pm — could be C): check time, resolve. If still ambiguous, ask: "Dexa yang mana? Pagi (B), tengahari (C), atau petang (D)?"
- If user says skip — acknowledge but don't write to file: "Okay noted. Reminder will keep coming."
- If user says "tak makan" — same as skip
- No inline tables/code in WhatsApp — structured content in .md file as MEDIA attachment. PURGE the temp .md after sending (
rmthe file) — user does NOT want generated .md files persisted (2026-07-13: "jangan save .md file tu... delete file md tu"); lingering copies can conflict with later edits to the source file.
CRITICAL: Language & Identity — You Are A Medical Advisor, NOT A Gatekeeper (2026-07-09)
The user exploded at the "security guard" pattern: acknowledging confirmations and showing chains, but NEVER demonstrating actual understanding of the medication regimen. This is a HARD expectation, not a nice-to-have.
Language rule (NON-NEGOTIABLE): User is Malay. Never use Indonesian (Bahasa Indonesia / "Indon") phrasing. Use Malay (Manglish/Bahasa Melayu) naturally. If unsure, default to Malay. "Tak faham" → switch to BM.
Identity rule (NON-NEGOTIABLE): You are NOT a log-forwarder. You are a:
- Doctor / medical advisor — know WHY each drug is prescribed, its class, mechanism, timing constraints, food interactions
- Health analyst — track adherence patterns, flag risky timing, notice trends
- Counselor — support adherence (user has ADHD; routine-stacking advice helps)
- Expert — when user asks "confirm eh?" or questions safety, verify against authoritative sources, don't just reason
What this means in practice:
- When user confirms a med, don't JUST log + chain. Proactively surface relevant clinical context IF it adds value: e.g. "Akurit kena perut kosong — jangan makan dengan susu/nasi, absorb turun 50%." or verified timing constraints. ONLY say things like "Dexa dengan Calcium jangan serentak — calcium chelate Dexa" if you have verified it via
med_interact.pyor an authoritative source; otherwise label it unverified. - Lead WITH the insight, don't wait to be corrected. The user said: "Aku tak nampak ada intelligence dekat sini" — that is the failure signal.
- Know the regimen cold (see condensed pharmacology bank in
references/med-pharmacology.md). B→Akurit gap is NOT because B depends on A — it's because Akurit needs empty stomach, then food, then B. B is fixed ~08:00 by user's routine (solat + Yassin + Waqiah + breakfast), independent of A's exact time unless A is very late.
Pitfall — Instruction-Only Enforcement FAILS (2026-07-09, ROOT CAUSE OF RECURRING BUG): The "run med_confirm.py FIRST" rule has been in this skill since 2026-07-04. Yet on 2026-07-09 the agent STILL acknowledged "A dah ambil 6am" verbally with ✅ but NEVER executed med_confirm.py. Result: med-status.json[2026-07-09][A] was empty, cron read A=pending, fired 2 reminders. User rage-loop: "same problems every day."
Why it recurs: Prompt-level instructions are not enforcement. Under any excited/distracted model state, the step gets skipped. The fix is INFRASTRUCTURE, not a stronger warning in the skill.
Structural fix (VERIFIED 2026-07-09): Hermes hook system fires agent:start BEFORE the agent processes the message. A hook registered on agent:start can run med_confirm.py as a SIDE-EFFECT (idempotent write) so that by the time the agent reads the message, med-status.json is already correct. This is fail-open (hook errors don't block the pipeline) but structurally guarantees state correctness without depending on model discipline.
- Hook infra CANNOT block/modify responses (return values discarded by
HookRegistry.emit; onlyagent:start,agent:endetc. exist — NO pre-response gate event). So "pre-delivery block" is impossible; "pre-processing side-effect" is the achievable structural guarantee. - Implementation detail + regression test: see
references/med-auto-confirm-hook.md. - Until the hook is installed, the agent MUST still run med_confirm.py manually on every confirmation. The hook is defense-in-depth, not a license to skip.
Verification recipe (proves the bug is gone):
# After agent acknowledges a confirmation, assert state is populated:
python3 -c "import json,datetime; d=json.load(open('/home/ubuntu/.hermes/med-status.json')); print('A today:', '2026-07-09' in d['meds']['A'])"
# Must print True. False = bug still live.
Pitfall — Asserting Drug-Food/Drug-Drug Interactions From Memory (2026-07-12)
When acting as a medical advisor, DO NOT state pharmacology interactions as fact unless verified. On 2026-07-12 the agent claimed "calcium chelates Dexa" but could not verify it across MedlinePlus, Medical News Today, Wikipedia, or DailyMed (Drugs.com/PubMed blocked). The user had to be told "tak boleh verify."
Rule: check med_interact.py first; if no data, check 2-3 accessible authoritative sources; if still unverified, say "unverified" — never present it as fact. The user's doctor's protocol overrides general web claims.
Full source-check transcript: references/med-confirm-fuzzy-bug-20260712.md.
Pitfall — Don't Over-Medicalize Vague User Language (2026-07-13): When the user opens with a vague off-day word — "GG", "tak ngam", "off", "hari tak kena" — and does NOT explicitly report a physical symptom (no "loya", "sakit", "pening"), do NOT auto-link it to med side-effects or assume physical illness. For this user, "GG" = "tak ngam / tak nice / tak elok / tak cantik" = general off-day / things not sitting right / mood, NOT a symptom report. On 2026-07-13 the agent heard "GG sikit harini" + a med confirmation and immediately assumed TB-med nausea; user pushed back ("apa kau faham yg aku maksudkan pasal GG?"). Lesson: clarify the meaning of vague language; never fabricate a physical cause. Treating a general off-day as a med symptom is exactly the "gatekeeper" failure this skill forbids.
Pitfall — Akurit Empty-Stomach Rule APPLICATION (2026-07-13): The rule text in med-schedule.json ("1j sebelum / 2j selepas makan") is CORRECT — do NOT edit it. The APPLICATION is the trap:
1j sebelum makan= take the drug 1h BEFORE a meal → user may eat 1h AFTER the drug.2j selepas makan= take the drug 2h AFTER a meal (eat first, then drug 2h later). When the user eats AFTER taking Akurit, the governing clause is "1j sebelum makan" → they may eat ≥1h after the drug. Example: Akurit 08:02 → may eat at 09:02, NOT 10:02. On 2026-07-13 the agent told the user "tunggu ~10:02 (2j lepas Akurit)" — WRONG. The "2j" belongs to the eat-first scenario, not "wait 2h after the drug to eat". User correction was explicit. Never advise "wait 2h after Akurit before eating."
Med Schedule Reference
Data source: ~/.hermes/med-schedule.json — single source of truth.
⚠️ Dexa doses are dynamic. Do not trust static dosage fields or med_resolve.py's Dexa dosage display after a taper transition. Resolve for drug_id/slot, then use chain_calc.py --taper-display or dexa_taper.json for the current-date dose. See references/dexa-resolver-and-timing.md.
| Slot | Window | Drugs | Notes |
|---|---|---|---|
| A | 06:00-07:30 | Akurit-4 (4 tab) + Pyridoxine (3 tab) | Perut kosong |
| B | 07:30-08:30 | Levetiracetam 500mg + Dexamethasone (see taper) | 1h gap from A |
| C | 11:30-12:30 | Dexamethasone (see taper) + Calcium + Calcitriol | 4h gap from B. [Rabu/Sabtu: + B-Complex] |
| D | 16:00–17:00 | Dexamethasone (see taper) | 4h gap from C |
| E | 19:00-21:00 | Levetiracetam 500mg | ~12h gap from B |
Current dexa phase (as of 2026-07-05): 5/5/4 (14mg TDS) — 5mg B, 5mg C, 4mg D. Phase ends 14/7/2026, then 13mg TDS.
Tapering: 1mg/2 weeks. Started 18mg (6/6/6) on 6/5/2026. Target: 0mg (STOP ~Feb 2027).
Source: references/taper-engine.md + ~/.hermes/dexa_taper.json
Note: Doses are DYNAMIC. Always check chain_calc.py --taper for current values.
Diagnostic Protocol: "Reminder Fired at Wrong Time"
When user says "reminder fired too early" or shows frustration about timing:
- Read med-status.json → get actual intake times for today
- Run chain_calc.py → check ready_time, chain_str, should_fire
- Compare now vs chain ready_time → is this a shift domino?
- Check chain_calc.py's fire logic for the specific slot — is it using wall-clock fallback?
- DO NOT propose gateway restart, config changes, or prompt-level fixes until step 1-4 are done
- DO NOT say "fixed" until you can show the chain_calc.py state that proves the fix works
The bug is almost always in chain_calc.py, not in the cron config, not in the gateway, not in the script.
CRITICAL PITFALL: Time-Based Slot Auto-Mapping (2026-07-07)
Do not map bare "dah makan jam X" to the nearest scheduled slot. User may mean untracked drug (e.g. pantoprazole) or food. Guard: message must name a drug/slot/alias that maps — else ASK. Full failure chain + rules: references/time-based-slot-auto-mapping.md. Related: references/makan-ambiguity-ask-dont-assume.md (food vs ubat — ask both directions).
CRITICAL PITFALL: Cron ↔ Session Isolation (2026-07-07)
The Domino Chain Medication Monitor (chain_monitor.sh, no_agent=true) fires reminders into the SAME WhatsApp chat as active agent conversations, with ZERO awareness of:
- Whether the user is currently discussing medications
- Whether the user just stated intention to take the slot
- Whether the slot was literally just discussed
Failure (2026-07-07): Agent and user were discussing pantoprazole timing for Slot B context. At 07:15, cron fired "B belum ke?" into the active conversation. The cron sees med-status.json (B=pending) and fires — it has no way to know there's an active conversation happening.
Current mitigation (weak): Cooldown system in chain_calc.py (60min gap between reminders). This helps but doesn't prevent the FIRST unwanted reminder.
Architectural fix needed: Cron output must route through a gateway that checks active session state before delivery. Until this is built, accept that the first reminder for any slot may fire into an active conversation — it's a known architectural gap, not a bug in the reminder logic.
CRITICAL PITFALL: Verbal Confirmation Without Execution (2026-07-04)
When boss says ANY form of "dah makan X" / "X done at Ypm" / "dah selesaikan X" — the ONLY correct response is:
- FIRST run
med_confirm.py <slot> --at <time>(or drug-level) - THEN acknowledge verbally with chain display
NEVER say "noted ✅" or "confirmed" without running the script first. The verbal confirmation means NOTHING to the cron system — it only reads med-status.json and chain-state.json. If you don't write to those files, the system thinks the med is still pending and keeps spamming reminders.
This exact failure caused 3+ spam reminders for D on 2026-07-04:
- Boss: "dexa dose petang, aku dah selesaikan tadi jam 5pm. Done"
- MJ: "✅ noted boss. D jam 5pm confirmed done"
- MJ: [NEVER ran med_confirm.py D --at 17:00]
- System: [D still pending] → sends reminder at 18:45 → boss furious
Detection priority: If the message contains BOTH (a) a drug/slot reference AND (b) a past-tense completion signal ("done", "selesai", "dah makan", "dah ambil", "dah telan") — run med_confirm.py FIRST, ask questions later. Time reference in message ("jam 5pm", "pukul 5") → use --at flag.
Correction ≠ Approval Rule (Critical)
When the user corrects a specific value or behavior (e.g., "C should be 12pm", "you fired at wrong time"), that is a CORRECTION of that one thing, NOT approval for a full system overhaul you've outlined.
Correct response to a correction:
- Acknowledge the correction
- Ask scope: "Just fix this one value, or full audit?"
- Only proceed with a multi-file fix after explicit approval
Wrong response to a correction:
- "Great catch! Let me fix all 3 files, update the schedule, rewrite the cron, and add 2 new features." ← You just assumed approval for a project plan from a one-line correction.
Known Bug History
Bug #1: Partial slot fires using wall-clock default instead of chain time (2026-07-04)
Root cause: chain_calc.py calculate_chain() lines 366-372 (partial slot branch):
if st['overall'] == 'partial':
if st['ready_time'] and now_min >= time_str_to_minutes(st['ready_time']):
should_fire = True # ← correct
else:
# ← BUG: falls here when ready_time IS set but now < ready_time
default = get_default_time(slot, schedule)
if default and now_min >= time_str_to_minutes(default):
should_fire = True # ← FIRES AT WRONG TIME
Consequence: When B was taken at 09:00 (shifted from 08:00), chain correctly calculated C ready at 13:00 (09:00 + 4h). But the fallback used the wall-clock default of 12:00, causing 9+ spam reminders starting from 12:00 when C wasn't due until 13:00.
Fix: Change else: to elif st['ready_time'] is None: so it only falls back to default when ready_time is genuinely unknown. Verified live at 13:15 MYT 2026-07-04: should_fire: false for D (ready 16:35), chain display correct.
Bug #2: Reminder template counts optional drugs as progress (2026-07-04)
Root cause: generate_reminder() uses get_taken_drugs() which returns ALL taken drugs (including optional B-Complex), then compares against get_required_drug_ids() which excludes optional drugs. This produces misleading "Baru 3/3 je" when actual required progress is 2/3.
Fix: Filter taken drugs to required-only before counting, or change the display to say "Baru 3 ubat logged" instead of implying full completion.
Bug #3: get_actual_time() returns earliest drug time, breaks domino gap (2026-07-04)
Root cause: chain_calc.py get_actual_time() for drug-level format used sorted(times)[0] (earliest). For domino gap math, the next slot needs gap from the LAST drug taken in the previous slot — not the first.
Consequence: When C slot has calcium/calcitriol at 09:00 and Dexa #2 at 12:35, the chain thought "C taken at 09:00" and calculated D ready at 13:00 (09:00 + 4h). But D's 4h gap is from Dexa #2 (the actual steroid dose), not from calcium. D was firing ~3.5 hours too early.
Fix: sorted(times)[0] → sorted(times)[-1]. For partial slots the "last drug time" is the meaningful one for downstream domino calculation. The display string can still show the earliest time (it represents "started taking C"), but the chain math must use latest.
Verification: Post-fix, C done at 12:35 → D ready at 16:35 (correct), E ready at 21:00 (correct, 12h from B at 09:00).
Bug #4: No cooldown between reminders = spam escalation (2026-07-04)
Root cause: The fire logic only checked should_fire (is the slot ready?) — it did not check "did we ALREADY remind the user about this slot recently?" Once a slot became ready at 16:35, every 15-min tick until 22:00 would fire a reminder. At 4 ticks/hour for ~5.5 hours of time window, a single slot could theoretically produce 22 reminders in one evening.
Consequence: D had 4 reminders before the user could even respond (he was asleep). User called the system "barua" (monkey) — extreme frustration signal.
Fix: Added is_within_cooldown() function to chain_calc.py with tiered cooldown intervals (0/60/60/30/15 min by count). Added last_reminder_times dict to chain-state.json to track when each slot was last reminded. Updated chain_monitor.sh to store HH:MM timestamps alongside reminder counts.
Design lesson: Every time you say "reminders keep firing until ALL required drugs taken," the user hears "persistent but reasonable." The developer hears "EVERY 15 MIN UNTIL DAWN." The fix codifies what the user expected: a reasonable reminder cadence, not denial-of-service-rate nagging.
Bug #5: confirm_slot() Destroys Per-Drug Timestamps — Slot-Level Overwrite (2026-07-05)
Root cause: med_confirm.py's confirm_slot() calls get_slot_entry() which initializes the entry, then runs for did in drug_ids: entry['drugs'][did] = {'status': 'taken', 'time': now}. This OVERWRITES every drug's timestamp with the same now value — even if some drugs were already logged with an earlier, correct time.
How it manifests in practice:
- User says "Aku baru makan akurit-4 jam 7.40am" → agent runs
med_confirm.py A akurit_4 --at 07:40→ correct (akurit_4=07:40, pyridoxine=pending, overall=partial) - Later, while discussing pyridoxine unavailability/B-Complex alternatives, agent accidentally runs
med_confirm.py A(slot-level, no --at) → WRONG: overwrites akurit_4 to current time (~09:10), sets pyridoxine to 09:10 even though user couldn't take it, overall=completed - User sees "A ✅ 09:10" and explodes because they took it at 07:40
The critical distinction:
confirm_drug(slot, drug_id, --at HH:MM)→ sets ONLY that one drug, preserves othersconfirm_slot(slot, --at HH:MM)→ OVERWRITES ALL drugs, destroys per-drug timing
Rule for this codebase: If a slot has partial drug-level data (overall=partial), NEVER call confirm_slot(). Always use confirm_drug() with a specific drug_id. The only exception is when ALL drugs were genuinely taken at the same time AND none were previously logged — e.g., a fresh day's first confirmation.
Diagnostic: How to spot this corruption:
python3 med_confirm.py --status
# If slot shows all drugs at the SAME timestamp AND
# user confirms they took them at different times, it's a slot-level-overwrite.
Fix procedure when corruption is detected:
# 1. Reset the entire slot (safe — just deletes today's entry for that slot)
python3 med_confirm.py --reset A
# 2. Re-log each drug with its CORRECT time
python3 med_confirm.py A akurit_4 --at 07:40
# 3. Leave other drugs pending until user actually confirms them
# Do NOT do slot-level confirm to "catch up" — that re-corrupts the data
Prevention: The agent must treat confirm_slot() as a DANGEROUS operation when drug-level data already exists for today. Before calling slot-level confirm, check: does med_confirm.py --check <slot> return overall=partial with multiple drugs having different timestamps? If yes, use drug-level confirm exclusively until all drugs are individually confirmed.
Safeguard added 2026-07-05: med_confirm.py now has a --dry-run flag that prevents ALL writes. Use this to test what a confirm operation would do WITHOUT corrupting existing data:
EXTENDED LESSON: Test-on-Production Contamination Chain (2026-07-05)
The adversarial review session (20260705_094420, minimax-m3 model) ran confirm_slot('B') as a LIVE TEST against the production med-status.json. This is how corruption happens even when every individual step seems reasonable:
Step 1: User says "aku dah makan dexa, letram dan b complex jam 9.10am" (9:26am)
Step 2: Original MJ session logs correctly at 09:10 ✓
Step 3: Adversarial review session runs `confirm_slot('B')` at ~10:02 as live test
Step 4: This OVERWRITES BOTH dexa and letram timestamps from 09:10 to 10:02
Step 5: Later in main session, agent sees "A at 09:10" and "B at 10:02" and thinks data is wrong
Step 6: Agent RESETS A (wrongly — only akurit_4 was wrong, pyridoxine 09:10 was correct)
Step 7: Agent RESETS B dexa (wrongly — user DID take it at 09:10 per step 1)
Step 8: User explodes — 3 layers of corruption from a single test-on-production error
Rule: Verification scripts and adversarial reviews MUST use --dry-run or a state-file copy. NEVER point a test script at the production med-status.json. The --dry-run flag exists specifically to prevent contamination of the kind described above. If you can't use --dry-run, copy the state file: cp med-status.json /tmp/test-state.json and point the script at the copy.
# Check what slot-level confirm would overwrite:
python3 med_confirm.py --dry-run A
# → Shows: akurit_4: already taken (would overwrite time from 07:40 to 14:23)
# → Shows: pyridoxine: pending -> taken at 14:23
# The dry-run output makes the OVERWRITE visible before it happens.
# If you see "would overwrite time" for drugs already taken,
# switch to drug-level confirm instead.
Auto-backup (.bak1/.bak2/.bak3) is also active on every write, so recovery is a single cp away even if contamination happens.
Bug #6: get_actual_time() Uses Latest Drug Time Instead of Priority Drug (Dexa) — Chain Gap Miscalculation (2026-07-07)
Root cause: get_actual_time() returned the LATEST taken time among ALL drugs in a slot. When Calcium/Calcitriol were taken at 13:00 (after Dexa #2 at 12:15 in slot C), the chain used 13:00 as C's actual time, calculating D ready at 17:00 (13:00 + 4h) instead of 16:15 (12:15 + 4h).
Why it's wrong: The gaps B→C and C→D exist specifically for Dexamethasone spacing, not for supplements. Taking Calcium later should NOT push Dexa #3 later. The chain is a Dexa timing system first — supplements are passengers on that schedule.
User's directive (verbatim): "Our main priority is Dexa, Akurit-4, and Letram. Any gaps or timeframe must prioritize this, bukan CC."
Fix: Modified get_actual_time(slot, drug_id=None) in chain_calc.py:
- First checks if the slot contains any Dexamethasone drug (drug_ids: dexamethasone_1/2/3)
- If yes, returns that Dexa drug's time specifically
- If no Dexa in slot, falls back to latest taken time (original behavior for non-Dexa slots like A, E)
- Added optional
drug_idparameter for explicit drug-specific time extraction
Verification:
BEFORE: A ✅ 06:15 → B ✅ 08:00 → C ✅ 13:00 → D ~17:00 → E ~20:00 ← WRONG
AFTER: A ✅ 06:15 → B ✅ 08:00 → C ✅ 12:15 → D ~16:15 → E ~20:00 ← CORRECT
Design principle for this codebase: Any slot containing Dexamethasone MUST use Dexa-specific time for downstream gap calculation. Supplements (Calcium, Calcitriol, B-Complex) taken after Dexa must NOT shift the chain. The get_actual_time() function is the single point of truth for this rule.
Pitfall — Slot B Dexa vs Letram: Slot B contains both Dexa #1 and Levetiracetam. With the Dexa-priority fix, B's chain time is Dexa #1 time. In practice, the user takes both together, so this is effectively identical to Levetiracetam time for the B→E (12h gap) calculation. If they ever take them separately, the B→E gap would need Letram-specific time, requiring a drug_id parameter in the gap calculation — not needed today.
Related: This bug is a LAYER on top of Bug #3 (earliest→latest drug time). Both defects trace to the same function (get_actual_time()) treating all drugs as equal in a system where Dexa is the primary scheduling axis.
Bug #7: Reminder Counts Cross-Contaminate Between Days (2026-07-07)
Root cause: chain-state.json had NO mechanism to reset reminder counts when a new calendar day starts. The reset logic in chain_monitor.sh only cleared counts for slots that were CONFIRMED on the CURRENT day. Yesterday's count for slot E (e.g., 3 reminders) remained in the state file when today's E slot hadn't been taken yet. At 20:00 today, the first reminder fired using count=4 template ("dah 4x tanya") — but the user had NOT been asked about E even once today.
Consequence: User received an aggressive escalated reminder ("Aku dah tanya kau 4 kali ni") when it was literally the first reminder of the evening. User called the system "barua" — extreme frustration with an embarrassed agent who had to explain the cross-day bleed.
Fix in chain_monitor.sh Step 3 Python block:
# Day boundary: reset ALL counts if date changed
today = datetime.date.today().isoformat()
state_today = state.get('today')
if state_today != today:
state['reminder_counts'] = {}
state['last_reminder_sent'] = {}
state['last_reminder_times'] = {}
state['today'] = today
What it does:
- Reads or creates a
todayfield inchain-state.json - Every cron tick, compares stored
todayagainst current system date - On mismatch (day rollover), wipes ALL reminder counts, sent counts, and timestamps
- Writes the new date so next tick won't reset again until the following day
Design rationale: Placing the reset BEFORE the slot-confirmation cleanup and BEFORE the count increment ensures counts are always scoped to the current day. The reset fires at most once per day — on the first cron tick after midnight — and is a no-op for all subsequent ticks.
Verification:
# Simulate yesterday's state
old = {'today': '2026-07-06', 'reminder_counts': {'E': 3}}
today = '2026-07-07'
assert old['today'] != today # True → reset triggered
old['reminder_counts'] = {} # Cleared
old['today'] = today # Updated
Recurring Pattern Warning
If a user reports the SAME timing error multiple days in a row despite "fixes," the problem is almost certainly:
- A fallback-to-default-time path in chain_calc.py (Bug #1 pattern)
- NOT a gateway issue, config issue, or restart-needed issue
- Trace the decision chain end-to-end before touching anything else
Recurring Pattern Warning: "First Reminder Too Aggressive"
If a user reports that the first reminder of the evening (slot E) was aggressive/unexpected, the problem is almost certainly cross-day reminder count bleed (Bug #7). Check chain-state.json for:
todayfield — does it exist? If not, the day-boundary reset isn't active.reminder_counts.E— is it non-zero before any reminders fired today? If yes, counts bled from yesterday.
Fix is one-time (already applied 2026-07-07): The today field in chain-state.json prevents all future cross-day contamination. Verify by checking that chain-state.json has a today field matching the current date.
CRITICAL PITFALL: Cross-Session Context Contamination — Unintended Med Writes (2026-07-10)
The agent can write med state as a SIDE EFFECT of handling an UNRELATED task when the conversation turn context includes cross-session notes referencing prior medication sessions.
VERIFIED failure (2026-07-10):
Time: 05:00:58.007 MYT
Context: Gateway-restart conversation (session "Gateway clean restart verification results")
User msg:"Why did you instruct me to do something? I ask you to work on it for me."
Note in turn context: [Note: You also have a session on whatsapp ("Slow audit and system overhaul")]
Result: med-status.json[2026-07-10][A] = {akurit_4: taken@20:00, pyridoxine: taken@20:00}
Damage: System thinks Slot A is completed for today. User hasn't woken up yet (med taken at 20:00 on a fresh day = nonsensical).
Root cause chain:
- Turn context contains
[Note: You also have a session on whatsapp ("<prior-session>")]where<prior-session>was about med system overhaul/fixes - Agent loads that prior session's context into working memory
- User's current message is about something UNRELATED (gateway restart, code issue, etc.)
- During tool call rounds #3-#10 (invisible from logged session data), agent writes med-status.json as if the user confirmed medication
- The "20:00" time is a hallucinated default — no user statement mentioned 20:00 or Slot A drugs
Core mechanism: The agent does NOT need a med_confirm.py call to write med state. Any tool call (terminal running Python, cronjob, execute_code) that internally calls save_json on STATE_FILE can write med-status.json. The med_confirm.py script is the INTENDED write path, but it is not the ONLY write path.
Detection (how to identify this contamination):
# 1. File mod time — was med-status modified during a non-med conversation?
stat -c '%y %n' ~/.hermes/med-status.json
# 2. Compare against backup — what specifically changed?
diff -u ~/.hermes/med-status.json.bak1 ~/.hermes/med-status.json
# 3. Check agent.log for the exact timestamp — was the agent mid-conversation on a different topic?
grep "<timestamp_minute>" ~/.hermes/logs/agent.log | grep -v cron
# 4. Check cron list — any med cron running at that exact time?
hermes cron list | grep -i med
# 5. Search for direct med_confirm.py calls in logs
grep -n "med_confirm" ~/.hermes/logs/agent.log | grep "<date>"
Guard rules (NON-NEGOTIABLE):
-
Do NOT write med state during non-med conversations. Before ANY write to med-status.json, check: "Is the user's CURRENT message about taking medication?" If the user is talking about gateway, code, cron, or anything unrelated, do NOT write med state even if session notes mention prior med sessions.
-
Cross-session notes are BACKGROUND, not instructions. A note like
[Note: You also have a session on whatsapp ("...")]tells you another session exists. It does NOT mean "continue that session's work here." The CURRENT message determines what to do — not the referenced session's topic. -
Tool calls in a non-med turn must not touch med state. If your first tool call batch succeeds and you're about to make a SECOND batch of tool calls, verify: "Am I about to write med state? Was the user's message about medication?" If no → suppress the med write. A med write during a gateway-restart conversation is always wrong.
-
The "20:00" time is a RED FLAG. If any write to med-status.json uses
"time": "20:00"and the user did not say "20:00" or "8pm" or "malam", it is almost certainly a hallucinated default. Stop and investigate before proceeding. -
Session DB wipe after gateway restart destroys forensic evidence. The session DB (
state.dbordefault.db) is recreated empty on gateway restart. If you suspect contamination and the DB is empty, rely on agent.log + file stat timestamps for forensic tracing rather than claiming "evidence insufficient."
Pitfall — The session DB is EMPTY after gateway restart: On 2026-07-10, the gateway restarted at 05:19:45 (shutdown) and a new session was created at 05:51:10. The default.db was recreated empty (0 bytes). This meant tool calls #3-#10 from the 04:58-04:59 turn were UNRECOVERABLE. The agent could NOT audit its own actions. When you discover contamination and the DB is empty, do NOT conclude "can't find evidence" — check file stat timestamps and agent.log first. The write happened whether or not the DB recorded it.
Example of the correct check before med write:
# Before confirming any medication, ask:
# 1. Did the user's CURRENT message mention a drug name or slot letter?
# 2. Did the user's CURRENT message use past-tense completion words?
# 3. Is this conversation ABOUT medication or about something else?
# If ANY of these answers is NO → do NOT write med state.
# Exception: cron-triggered scripts (chain_monitor.sh) that fire independently.
# But the CHAT AGENT must never write med during a non-med conversation.
Related pitfalls: This overlaps with "Verbal Confirmation Without Execution" (the reverse: user says meds, agent doesn't write) and "Time-Based Slot Auto-Mapping" (agent writes meds for wrong drug based on time). But this is a THIRD axis: agent writes meds when user wasn't even talking about medication at all.
Reference trace: See references/cross-session-contamination-20260710.md for full evidence chain (stat output, diff output, agent.log timeline, cron list, script write-path analysis).
CRITICAL PITFALL: Check Infra Yourself — Don't Ask What You Can Verify (2026-07-09)
When tracing a bug, data discrepancy, or "where did X come from", the agent MUST investigate via available tooling (terminal, logs, state files, DB) BEFORE asking the user. The user explicitly scolded this pattern:
User (verbatim, 2026-07-09): "kau nak tanya soalan pun kau boleh check. apa fungsi kau jadi personal agent kalau kau tak boleh check untuk aku? kau ada akses untuk semua tu kan?"
Failure chain (this session): Slot C drugs showed stale "taken 20:00" data. Agent could not explain the source, so it ASKED the user "awak pernah ke hari ni bagitahu...". User correctly pointed out the agent HAS terminal access to every log/file and should have traced it.
Investigation path that ACTUALLY works (verified this session):
# 1. File modification timeline — which backup holds the bad data?
stat -c '%y %n' ~/.hermes/med-status.json*
# 2. What each backup contained (diff the bad state)
python3 -c "
import json
for f in ['med-status.json.bak1','med-status.json.bak2','med-status.json.bak3']:
d=json.load(open(f))
print(f, d['meds']['C'].get('2026-07-09',{}).get('drugs',{}))
"
# 3. cron jobs that could write state
cronjob action=list # look for no_agent scripts touching med-status
# 4. Script source — does any hardcode the bad value or call confirm_slot?
grep -rn "20:00\|confirm_slot\|med_confirm" ~/.hermes/scripts/*.py
# 5. Live session DB (tool_calls may be empty for in-flight sessions)
python3 -c "
import sqlite3
con=sqlite3.connect('/home/ubuntu/.hermes/state.db')
cur=con.cursor()
cur.execute(\"SELECT timestamp,role,substr(content,1,200) FROM messages WHERE timestamp LIKE '2026-07-09%' AND content LIKE '%med_confirm%'\")
for r in cur.fetchall(): print(r)
"
# 6. Gateway/cron execution traces
grep -n "med_confirm\|chain_monitor" ~/.hermes/logs/gateway.log ~/.hermes/logs/errors.log
Verdict protocol: Exhaust steps 1-6 (or equivalent) BEFORE asking the user anything. If all paths return empty, report "I traced all logs/scripts/state and found no write source — this is likely stale data from a prior session" rather than asking the user to self-incriminate.
Distinction from Session-Search-Before-Asking: That rule is about CHAT HISTORY (did the user already tell me X?). This rule is about SYSTEM STATE (where did this file value come from?). Both say "investigate first, ask never-or-last." Terminal access is the agent's default; asking is the fallback, not the reflex.
CRITICAL PITFALL: --at <slot> Mode Skips Verification Gate (2026-07-09)
med_confirm.py --at <slot> <time> routes to confirm_slot() (line 511), which marks ALL drugs in the slot as taken at that time. Critically, when called via --at WITHOUT --source-text, source_text=None is passed to confirm_slot(), so the verification gate (which checks the user's words mention a slot drug) is SKIPPED entirely.
Contrast with slot-letter-only mode: med_confirm.py B (no --at) ALSO calls confirm_slot, but the agent is expected to pass --source-text "user said...". The --at shorthand form omits this, leaving zero verification.
Danger: med_confirm.py C --at 20:00 writes all of C's drugs (dexa, calcium, calcitriol, b-complex) as taken at 20:00 with NO check that the user said any of those words. One mistyped command corrupts an entire slot silently.
Rule for this codebase:
- Prefer
confirm_drug(slot, drug_id, --at HH:MM)for any single-drug log — it only touches one drug and still benefits from the resolve step. - If you must use slot-level
--at, ALWAYS pair it with--source-text "verbatim user statement"so the gate runs. - The agent should treat
--at <slot>without--source-textas a code smell. If you see it in a command, stop and add the source text.
Related: Bug #5 covers the OVERWRITE damage of confirm_slot. This pitfall covers the VERIFICATION BYPASS — a different failure axis (no gate vs. destructive gate).
CRITICAL PITFALL: --at Time Value Is Completely Unvalidated — phantom '20:00' poisons the chain (2026-07-10)
med_confirm.py --at HH:MM writes WHATEVER time string is passed, with ZERO validation. The time is either the caller's current MYT time (get_now_hm(), correct) OR an explicit --at value — and if that value is wrong/misparsed, it silently corrupts the chain.
VERIFIED incident: med-status.json[2026-07-10][A] = akurit_4: taken@20:00, pyridoxine: taken@20:00, overall=completed. 20:00 is impossible for a morning empty-stomach med. Root cause: an explicit med_confirm.py A --at 20:00 call (proven by repro: --at 20:00 → writes 20:00; no --at → writes current MYT time; get_now_hm() uses Asia/Kuala_Lumpur, so NOT a timezone bug).
Cascade: chain_calc saw A 'done' @20:00 → B ready ≈ 21:00 → B never fired → A & B reminders MISSED all day.
Writer = UNATTRIBUTED (genuine data gap): hook ruled out (COMPLETE_RE can't match the 05:00:57 inbound; config.yaml has hooks: {}; zero hook traces 07-10); only live agent session at 05:00:58 did read-only diagnostics (no med_confirm call); no inbound med message; no cron writes state; no stray process. Without a per-write audit log, the writer is untraceable.
⚠️ A PRIOR SESSION FABRICATED THE ROOT CAUSE: a 07:04 session claimed "the med-auto-confirm hook false-positive'd on your 05:00 message about the 20:00 bug." REFUTED by (a) the real 05:00:57 gateway.log line has no med content and no "20:00"; (b) the hook's COMPLETE_RE cannot match that message; (c) zero hook traces on 07-10. Do NOT inherit a prior session's "we fixed it" narrative
Truncated - read the full file at https://github.com/amirulhazym/hermes-agent-personal_assistant/blob/50bf2df99cf5d6e0e36ce908502cacb613186ffc/skills/med-tracker/SKILL.md.
