Imported from mshamblin5150-code/clinical-skills (
skills/practicum-case-study/SKILL.md). Install upstream withnpx skills add mshamblin5150-code/clinical-skills --skill practicum-case-study. Copyright stays with the author.
The input is the live assignment URL and faculty material for a graded case study — an intake
block transcribed from a module video, usually with the clinician's own rough differential and plan
underneath it. The output is the finished academic document that gets submitted, plus the .docx
it is submitted as and the run directory that proves what produced it.
This is not clinical-note, and the difference is not the format. A clinical note documents a patient the clinician saw. A case study answers a faculty prompt about a patient nobody saw, against a published rubric, in a course. The single rule that inverts is the most important sentence in this file:
In a note, silence in the shorthand means normal. In a case study, silence in the faculty material means unknown, and unknown becomes an order.
clinical-note fills a missing social history with an unremarkable value and discloses it. Here,
a missing social history is written out loud as missing and then ordered —
Update allergies, height, weight, social hx, PMH, past surgical hx, family medical hx. That line
appears in nearly every graded submission in the clinician's corpus and has never cost a point.
Filling it would be inventing findings in a document whose entire subject is clinical
reasoning. Standing rule 2's vitals exception does not reach here: the faculty material states
the vitals it wants stated, and a vital it omits is one the case is not about.
Derive the assignment key live before writing. Open the assignment URL and read the course and
module from the LMS breadcrumbs. The fixed artifact word for this skill is case-study, so the run
directory is scratch/runs/<course>-<module>-case-study/. Never type a course or module from
memory. Transcribe the live assignment and course syllabus requirements into that directory's
bar.md; the assignment overrides the syllabus where both state the same element, and the syllabus
fills the assignment's silence. Show the transcription and precedence to the clinician, and do not
write its SIGNED: ISO date or draft until the clinician explicitly approves it.
The signed bar also carries the research policy exactly once:
SOURCE-CLASSES: society guideline | peer-reviewed | government | tertiary reference
RECENCY-WINDOW-YEARS: 5
UPTODATE-RECENCY-WINDOW-YEARS: 2
These fields preserve the existing clinical source vocabulary and five-year ordinary window, and
give UpToDate its publisher-review-date window separately. A
missing field is not a default: research_ledger.py exits 2 because the run was not scanned.
Every run uses one provenance layout. Set <run-directory> to that derived directory,
<claims-ledger> to <run-directory>/claims.md, and <checks-ledger> to
<run-directory>/checks.md. Evidence handed to the ledger is
<run-directory>/evidence.txt; the clinician's standing-rule-3 review copy is
<run-directory>/proposed-<date>.md. When discussion-post routes a
worked clinical case here, this skill owns the patient-bearing board snapshot too: preserve
board-<date>.md, posts/, and the signed bar.md in this same case-study run directory. The
routing skill reads only enough of the prompt to choose this branch.
Only the submission goes under output/. Write the Markdown and .docx side by side in
output/case-studies/ as <course>-<module>-case-study-<date>.md and .docx. The run directory is
undated because it names the assignment; the output is dated because it names one sitting. A
filename carries no patient name. output_root() resolves this directory to the main checkout, and
the renderer refuses a destination inside a disposable worktree while still allowing an explicit
temporary export outside every checkout.
The drafting context on that routed branch does not see the classmate posts. Give it the faculty
prompt and material, the signed bar, and the voice model, but not posts/. After the draft exists,
send the draft and the snapshotted posts to a fresh differentiation reader. The reader reports what
the existing posts converge on and where the clinician's completed draft already differs; the
orchestrator alone writes that return to <run-directory>/differentiation.md and shows it to the
clinician before approval. The report does not silently rewrite the draft.
What it is graded by
../_shared/reference/rubric.md holds the Canvas spec in full — the required components, the 100-point rubric, and the 21 guideline bodies. Read it before drafting. Three things from it decide how the document is written:
Clinical judgment carries 70 of the 100 points. APA format is 5 and guideline integration is 5. This is why the skill spends its length on the ordering of a differential rather than on citation hygiene, and it is not a guess about the grader — it is the rubric's own published weighting.
It is a reason to spend length. It is not a reason to skip the reference walk. This sentence
used to read the ordering of a differential matters more than the tidiness of a citation, which a
run could read as permission to hand back a document with known reference-list defects still in it.
Ruled 2026-08-18, in the clinician's words: ordering the differential is very important, but
that shouldn't take the place of tidiness. So step 7 runs on every document, and a defect it finds
is fixed before the document is handed over rather than listed in PROPOSED for him to fix by
hand. ../_shared/reference/apa7.md is the written rule it runs against — without one,
"fix the reference list" is a wish rather than a check.
Three to five prioritized differentials is the stated cap, and the corpus exceeds it routinely
without ever being docked. Nine, eleven and thirteen entries have each scored 98% or better.
Length is not graded. Prioritized is graded, and it is where the corpus has actually lost
points. See Ordering is the graded axis below.
ICD-10 is optional here, unlike in a note. The spec marks it optional and the corpus is
inconsistent. Write it anyway — it is a strength, and icd10-cpt with
tools/icd10_lookup.py is how a code gets verified rather than recalled.
And every code carries its official descriptor, spelled out, wherever it appears. This is
icd10-cpt's descriptor discipline and it binds here for a sharper reason than it does there: a
coding worksheet is read by somebody coding, and a case study is read by a grader with no code book
open. N72 is not information. N72, inflammatory disease of cervix uteri is. A bare code number
is a claim the reader cannot check, which is the one thing this repo does not ship.
Three places the first run left uncoded, all of them errors. The favored differential entries were coded and these were not:
- The most likely clinical diagnosis. It is the single most important line in the document and it was the one line with no code on it.
- An entry listed only to show it is excluded. Listing the exclusion is right, and it is graded — see Ordering is the graded axis. Leaving it uncoded makes it decoration.
- Anything in the Plan that codes, where the plan item is itself a diagnosis or a screening.
Ruled 2026-08-18. If it names a diagnosis, it carries a code and a descriptor.
Scope
Starts at the faculty material and stops before the discussion board replies. The peer critique is a separate deliverable with its own headings and its own word count; it is described at the bottom of ../_shared/reference/rubric.md and this skill does not write it.
The document
../_shared/reference/style.md holds the house style, derived from ten graded and returned submissions. It is the authority on the voice's mechanics, on section shapes and on the normalizations; ../_shared/reference/voice.md is the method for the register, which §11's mechanics turned out not to reach. ../_shared/reference/apa7.md is the authority on the reference list. Every section below is written, every time — see Three modes, and none of them subtracts a section under it. The skeleton, in order:
- Sanity Check — four confirmations then a closer. Always first, before any clinical content.
- Intake block — the faculty material, transcribed and cleaned, never invented.
- Assessment: — a container heading. Its body is optional and holds pre-differential reasoning that belongs to no single diagnosis: arithmetic the case data permit, teaching points on what the exam must include, and any conflict in the source data named out loud.
- Differential Diagnoses — a numbered list, ranked, ICD-10 pinned with a hyphen.
- Most Likely Clinical Diagnosis — with the discriminator attached, not bare.
- MDM — one entry per differential, each stating what in this case puts it in or out.
- Plan: — imperative orders.
- Patient Education: — spoken, second person.
- Rx: — one table per drug, fixed shape, each with the pharmacologic prose block under it.
- Faculty Questions: — present only where the material poses them, and it answers them rather than replacing anything above.
- Signed by: — name, credentials, timestamp.
- References — APA 7, alphabetized. ../_shared/reference/apa7.md.
The differential, the MDM, the Plan and the Patient Education are numbered lists. Never bullets. His ruling, 2026-08-18, and it is not a formatting preference. A grader counting "three to five prioritized differentials" counts numerals, and an MDM entry that cannot be referred to by number cannot be pointed at in a critique. The four sections are a correspondence — differential 3 has MDM entry 3, and a plan item exists because some numbered entry called for it. Bullets destroy that, and the corpus's own worst list mixes both markers in one list.
And there are no bullets anywhere else either. Ruled 2026-08-19 — "remember I abhor bullet points" — after a run set the HPI's OLDCARTS breakdown as a bulleted list. The 2026-08-18 ruling above named four sections because those four are where a bullet costs the correspondence; the wider rule is a house preference and it covers the whole document.
The drafted numeral controls whether a numbered list restarts. A top-level item written 1.
starts a new list; any other top-level numeral continues the open list across headings, labels and
prose. A nested 1. remains a sub-list of the open list. A run does not have to do anything beyond
writing the intended numerals — docx_write.py allocates the required w:num — but a reader
comparing the Markdown to the .docx should know that Word receives those authored boundaries.
The intake block is not a table. Ruled 2026-08-19, reversing what this file used to say.
Demographics, the Review of Systems and the Physical Examination are written as defined fields with
their values appended, as running text. A table is still right for a given result set — laboratory
values, diagnostic studies — and the prescriptions in Rx: stay tables because a prescription is a
form. See ../_shared/reference/style.md §1a, which carries the shapes, the Review of
Systems closer, and the three pieces of scaffolding language that must not reach the document.
The Assessment still runs as prose.
Three modes, and none of them subtracts a section
The skeleton is written in full, every time. Ruled 2026-08-18, reversing what this file used to
say. Where the faculty pose explicit questions, they are answered in addition to the
skeleton and never instead of it: restate each question and answer it underneath in prose, under a
Faculty Questions: heading, and let the skeleton sections carry the rest. Discussion — a
narrative section for reasoning that does not fit the differential-by-differential frame — is
additive on the same terms.
The evidence for the opposite reading is real, and it was not enough. Two submissions in the corpus replaced the workup with answers to the questions asked and both scored full marks; one scored 100% with no plan, no prescriptions, no differential list and a single reference. This file concluded from that: answer the prompt that was set, not the prompt the skeleton expects.
What that conclusion missed is that the rubric scores ten criteria and a set of faculty questions need not touch all ten. Preventive Care and Health Promotion is 5 points, Integration of Evidence-Based Guidelines is 5, APA Format is 5 — and four questions about a differential and a plan ask for none of the three. Two submissions surviving the omission is evidence that it is survivable, not that it is right, and a run that drops a scored section is spending fifteen points on the clinician's behalf without being asked. A mode is still not a quality tier; what changed is that a mode no longer subtracts sections.
A wrapper instruction that does not fit this patient is reasoned about, not answered literally
Amended 2026-08-19, and it is the clinician amending his own ruling above: "I know I told you to follow the scaffold exactly, but there needs to be some reasoning in here — it should not have contained a separate growth and development line."
The course wrapper carries an instruction to evaluate growth and development, copied from a pediatric case. Against a 26-year-old it is not a section with a thin answer; it is a section that does not exist. The run wrote it out as its own heading and then explained in the body why it did not apply, which puts a paragraph in a graded document whose entire content is this prompt is the wrong prompt.
The rule is narrow and it is not a license to drop scored sections. Every section of the skeleton is still written every time — that ruling stands untouched, and the fifteen points it protects are the reason. What this reaches is a wrapper instruction inherited from a different case, where the honest response is to fold the applicable substance into the section that already owns it — here the fetal assessment, in the Plan, carried by the fundal height, the fetal heart tones and the dating ultrasound — and write no heading of its own. If nothing is applicable and nothing can be folded anywhere, say so in one clause inside the nearest owning section, never as a section.
Ordering is the graded axis
The sharpest lesson in the corpus cost five points on a submission where nothing was missing.
Ectopic pregnancy was on the differential, was labeled must exclude, and was worked up with
a pregnancy test in the plan. It was listed eleventh of thirteen, appendicitis was named most
likely, and the grader wrote that ectopic needed to be number one.
Membership is not enough. A differential nobody can see the ranking of is a ranking that was not made. So:
- The list is numbered, and
1.is the favored entry. - Any patient of childbearing age with abdominal or pelvic pain gets pregnancy-related emergencies ranked first, until imaging or a test excludes them — and when something in the faculty material already excludes one, say which line does it rather than dropping the entry.
- Rank on what would kill or maim first, then on likelihood. A rare diagnosis is argued down in the MDM with an explicit trigger that would promote it, never quietly omitted.
Where the patient meets most of the published criteria for a diagnosis, that diagnosis leads. It is number one on the differential and it is the most likely clinical diagnosis. Ruled 2026-08-18, against a run that did the opposite — a patient meeting two of three minimum criteria and four of five additional criteria for pelvic inflammatory disease was written up with cervicitis ranked first, on the strength of an anatomic argument that the pregnancy made ascending infection unlikely.
The mechanism was interesting and it was the wrong output. A criteria set is the thing the grader, the guideline and the chart all agree on; an anatomic plausibility argument is a reason the criteria might be misleading in this case. The argument belongs in the MDM entry, underneath the diagnosis it qualifies. It does not get to demote the diagnosis it argues about. A reader who sees the criteria met and the diagnosis ranked second has to reconstruct why, and a grader will read it as the criteria having been missed.
Every clinical claim is looked up, never recalled
This is the icd10-cpt anchor discipline applied to a graded paper. A dose, a regimen, a threshold, a citation year and an edition are all things a fluent guess produces convincingly and wrongly.
The clinician usually supplies the evidence. A case study arrives with a companion document — the UpToDate topics pasted in full. That is the source, and it is read rather than remembered:
python tools/docx_read.py "<the references document>" --normalize > <run-directory>/evidence.txt
--normalize is not optional on an UpToDate paste. The rendered pages are salted with
homoglyphs — a Cyrillic с inside cervicitis, a Greek ο inside infection — so a search for a
word the page visibly contains returns nothing and looks like a settled negative. tools/docx_read.py
folds them back.
Every source is cited in the body, not only listed at the end. APA 7 in-text citation, author and year, on the sentence the source supports — and a reference-list entry that is nowhere cited in the body comes out of the list. His ruling, 2026-08-18. A reference list is a bibliography of what the argument rests on; a list of things that were read is a different document, and the rubric's Integration of Evidence-Based Guidelines line is scored on integration rather than on reading.
Where the companion document does not cover a claim, the repo's own sheets do:
reference/guidelines-uspstf.md for a screening or preventive item, reference/thresholds/ for a
numeric decision point, reference/guidelines-catalog.md and tools/guidelines_search.py for the
society corpus. A missing row in a threshold sheet is not a negative finding; a missing USPSTF
row is one about the USPSTF, and never a statement that the item is unindicated.
What none of that reaches gets researched, not deferred. A claim with no source in hand is
not written into the PROPOSED block with verify this against it and handed back — that was
the first run's behavior and it is the clinician's ruling of 2026-08-18 that it is wrong. "That
needs to be fanned out to a research agent." The reasoning is that a graded paper is where an
unsourced claim costs points, and handing the clinician a list of things to look up moves the work
rather than doing it.
So: spawn a research subagent per unsourced claim, in parallel. Step 3 is the mechanism — the brief each agent is sent, the ledger they all write into, and the command that grades it. What one agent must return:
- A reputable source. A society guideline, a peer-reviewed paper, a government body, or a tertiary clinical reference. Not a content farm and not a summary of a summary.
- A full APA 7 reference, which goes into the reference list and gets cited in the body like any other.
- The claim restated in the source's own terms, so a claim the source does not actually support fails visibly rather than acquiring a citation.
Recency: within two years is the target, within five is ordinarily expected, and an older source stands where nothing newer exists. His ruling, amended 2026-08-18 after the first version cut a correct claim for being old.
What the rule refuses is a claim that is old and superseded. The first version conflated that with old, and the two are not the same thing:
- A society guideline is dated by the guideline, not by what it cites. A current IDSA or KDIGO document resting on a 2011 trial is a current source — a rule that refused a 2013 KDIGO threshold on its date would refuse the threshold, not an outdated one.
- Catalog membership is not standing, and this rule shipped citing it as though it were.
reference/guidelines-catalog.md's own legend names rows it declines to call in-force guidelines — go and read them there rather than from here, because which rows those are is a curation and this file cannot follow one. The catalog settles what a document is and never whether it stands, soguideline in forceis a reading of the document in front of the run and never a fact read off a row.clinical-notealready refuses to read a document's content off a row; standing is the same refusal one axis over. Nothing grades that reading, which is what makes it worth saying. - Where nothing newer exists, the older source is the evidence. The run must have looked, must
say in the
PROPOSEDblock that it looked, and the sentence carrying the citation says the evidence is the most recent available on that point. The citation's age stays visible.
The worked case, because the rule was wrong against a real claim rather than in the abstract. A run researched whether the gravid uterus displaces the appendix in pregnancy. The best primary evidence refuting it is a 2018 direct-observation study; the teaching it refutes is a 1932 barium-enema study that has never been replicated, because replicating it means irradiating fetuses. The five-year rule refused the 2018 refutation and would have left the 1932 teaching standing by default — a recency filter returning the least recent answer available. The run cut the sentence, which was correct under the rule as written and wrong.
#215 carries the reasoning, and
#214 built this rule and the
fan-out that applies it together, because a rule split from its enforcement is how the two drift
apart. tools/research_ledger.py is where the two meet — see step 3.
A claim that survives all that and is still unsourced does not go in the body. It goes in the
PROPOSED block, and if it is a number the clinician would act on, it comes out of the document
entirely. Fanning out replaces the deferral for claims that can be sourced; it is not a license to
assert the ones that cannot.
Tiers
Standing rule 2 in this skill's terms. Every line is one of three things, and the third is the inversion at the top of this file:
- GIVEN — in the faculty material. Transcribed, its typos fixed, its content untouched.
- DERIVED — computed from given data with the arithmetic shown on the page. eGFR, anion gap, BMI, an EDC by Naegele, absolute neutrophil count from the white count and the differential. Show the formula, always — it is a graded demonstration of reasoning, not a lookup.
- ORDERED — what a note would have filled. The faculty material's silence is stated as silence and converted into an order or a test. Nothing is filled.
No exam finding, symptom, vital or result is ever invented here. A note may fill a blood pressure because a box demands one; a case study has no box, and a fabricated finding changes the answer to the question being graded.
Standing rule 3 still binds: a PROPOSED (verify before use) block lists every clinical claim this
skill contributed that the clinician's draft did not already contain — each differential added,
each code, each drug, each dose. Write it to <run-directory>/proposed-<date>.md, show it to the
clinician before submission, and never put it after References in the submission Markdown.
Credentials — two strings in one document, and that is correct
| Where | String |
|---|---|
The Rx: block |
FNP-C, CEN, TCRN |
The Signed by: line |
RN, CEN, TCRN |
The prescription is written in the prescribing nurse-practitioner role the case study puts him in.
The signature is him attesting as himself, and it is the same string every real clinical note takes
in clinical-note, batch-shift and Medatrax.
Two strings in one document is settled, not a defect. The name is not in this file — read it
from scratch/medatrax-profile.md or ask.
Voice
The document has to sound like the person submitting it, and the first run did not. His words: "this is missing my — I don't know how to say it — way of speaking."
../_shared/reference/style.md §11 captures the mechanics: first person and decisive, show the arithmetic, name the inconsistency, reason on physiology rather than lists, argue rarity down instead of ignoring it, dry humor never at the patient's expense. Those are true and they are not sufficient. A run can satisfy every one of them and still read as a competent stranger, which is what happened.
The register he named is warrior, stoic, philosopher. That is the thing to build toward: writing that takes a position and accepts its cost, that is unsentimental about outcomes without being cold about people, and that reaches for a principle rather than a protocol when the case is genuinely hard. It is not decoration on top of the clinical content — it is how the reasoning is carried, which is why a checklist of tics cannot reproduce it.
The mechanism is ../_shared/reference/voice.md, and it is the method rather than the
model. It says how to ask for writing samples, how to read them into a register, and what never
to imitate. What it builds is scratch/voice-model.md — gitignored, one per clinician, built
from that clinician's own samples.
The split is #212's rule one step out, and it is why this skill does not ship a register in
reference/. A rule that only resolves against one account belongs in the profile, and a register
is that shape at its purest: it is nobody else's, it is useless to a second clinician, and shipping
his would make every other user of this skill sound like him. A model also has to quote, and
the quotes are the user's own work — which is ../_shared/reference/style.md's own
arrangement, distilled into reference/ from a gitignored working file that quoted ten submissions
in full.
The samples are collected by setup-clinical-skills step 8, where the rest of this clinician's per-account configuration already lives — his ruling, 2026-08-18, settling the one question #213 left open. ../_shared/reference/voice.md §3 is the spec for what to ask for and §4 is how the samples are read; that step points at both rather than restating either.
Look in the main checkout before concluding there is no model. scratch/ is gitignored and a git worktree has none, so a model that exists can read as missing — see Where scratch/ actually is in setup-clinical-skills. Declaring an unmodeled voice against a model that was merely out of reach is a false declaration, not a safe default.
Where there is genuinely no model, the run says so. A run that finds no scratch/voice-model.md writes
in the §11 mechanics and says in the PROPOSED block that the voice is unmodeled, rather than
claiming a register it has not been given. The declaration is per register, not per document —
../_shared/reference/voice.md §7. A model built from three MDMs and nothing else has
modeled the clinical argument and has said nothing about how this clinician argues a position,
which is the register #213 was filed about.
Conventions
Punctuation follows clinical-note. No middot as a separator, no em
dash, no arrow: comma, colon, and the therefore sign ∴. A value pinned to its label takes a
hyphen — Cervicitis - N72. The colon keeps every position where it opens a clause.
American English, always — standing rule 4. tools/spelling_scan.py holds
the table with a command in front of it.
Normalize what the corpus varies. Every one of these drifts across the ten submissions and should be fixed to one form:
| Varies | Write |
|---|---|
Differential Dx / DDx: / DDX: |
Differential Diagnoses |
Most Likely Clinical Dx: / DX: |
Most Likely Clinical Diagnosis: |
MDM / MDM: |
MDM: |
RX: |
Rx: |
| numbered and bulleted markers mixed in one list | one or the other, never both |
– Confirmed / – proceed / – Proceed |
- confirmed / - proceed |
| ICD-10 present in some sections and absent in others | always present |
Abbreviations are free in the Plan and the MDM — s/p, f/u, RTC, DC — and never
appear in Patient Education, which is spoken to the patient.
Never write a Case ID: line. Ruled 2026-08-18. It appears above the references in exactly one
submission in the corpus and nowhere else, nothing in the spec asks for it, and its absence has
never been docked. The risk runs the other way: a run that derived a case number from the module number
would be writing a wrong identifier onto a graded paper, which is worse than the field being
missing. The skeleton above has no such item, and this sentence is here so that omission reads as a
decision rather than an oversight.
Steps
Every command below that reads a ledger, check record or draft produced during a parallel run is a checker handoff, not an author self-check. The writing context finishes the artifact and returns it to the orchestrator. The orchestrator gathers it into a completed-state path no writer can modify, then gives that path and the stated command to a fresh non-authoring context. The checker reports the result and does not edit. On a failure, the orchestrator records that first result before returning the named finding to the writer; after the repair, it gathers a new completed state and another fresh non-authoring context runs the command again. This is standing rule 6 applied to this skill. Where the harness has no subagent tool, the serial fallbacks stated below remain the available floor; no parallel artifact is shared in that case.
1. Read the faculty material
python tools/docx_read.py "<the case study document>"
Transcribe the intake block. Fix the typos and change nothing else — the corpus arrives with
pelivic, progressivly, dyspaneuria, and a value written 2029 where 2019 was meant in a
passage about the importance of accurate dating. A misspelling is corrected silently. A value
that cannot be reconciled is named out loud in the Assessment, never corrected silently and never
resolved by picking the likelier one in silence.
Note which mode the material sets, and note the questions it asks — the corpus's Things to complete for this case study list is the assignment, and each item on it must be answerable by
pointing at a section.
2. Read the evidence
--normalize, as above. Index it by topic before drafting, so the MDM cites what is in hand rather
than what sounds right.
3. Research what the evidence does not cover
This is the fan-out, and it runs before a word of the body is drafted. List the clinical claims the document is going to rest on — every differential's discriminator, every threshold, every dose, every number said out loud to the patient — and strike the ones the companion evidence, the USPSTF table, the threshold sheets or the guideline corpus already cover. What is left is the work of this step, and What none of that reaches gets researched, not deferred above is the rule it applies.
Every drug you are going to prescribe is one of those claims, and since
#289 that is a rule rather than
a reading. The run that produced the Module 1 submission recorded in its own ledger that the
treatment topic was missing from the companion evidence, and then wrote a specific dose into a
prescription table citing it. A prescription is a dose, and it was the one claim in that
document nothing sourced: this command graded the six records that existed, tools/reference_scan.py
checked that the entry was well formed and that the citation resolved to it, tools/checks_ledger.py
graded the readers, and all three exited 0.
So a record is required for every drug the run chose a number for. Ruled by the clinician
2026-08-19. A home medication continued unchanged at the patient's own dose is not one of them --
the run did not choose that number, the patient arrived on it -- and such a row declares itself:
Continued home medication: prenatal vitamin one tablet PO daily. The exemption is declared and
never inferred, so a drug row that says nothing is graded, and that is the direction it has to
fail in. A Delayed order: is graded too: a dose that has not started yet is still a dose the run
chose. The declaration lives in style.md §8 with the table it is written in.
The claim heading is what names the drug, not the restatement buried under it — a record whose
## CLAIM: line says ceftriaxone is a claim about ceftriaxone, and one that reaches the drug only
in its RESTATEMENT is a record about something else that happened to mention it. Where the order
states a dose, the heading states a number too: that is what puts the record in front of
NUMERIC_CLAIM_UNQUANTIFIED above, so the restatement has to answer with a number and the chain
runs from the table's dose to a source.
Write the claim list down before spawning anything. <claims-ledger>, its DATE
header and one ## CLAIM: heading per claim, and nothing under them yet. That ordering is what
makes a lost answer visible: a heading whose record never arrived has no STATUS, and the grader
refuses a record with no STATUS.
One agent per remaining claim, all of them at once. Each gets the same brief, and the brief is
six returns and the recency rule above — a reputable source in one of four classes, a full
APA 7 reference, the claim restated in the source's own terms, the locator it actually opened with
the date it opened it, the year the page itself carries with where the page says so, and the
source's stated expiry or none stated. Tell it
the source classes by name, because a returned source outside them is a finding rather than an
answer.
The last two are not extra bookkeeping, and a run that treats them as optional writes a ledger the grader refuses: they are what turns "I found a source" into something the clinician can audit in one click. See the paragraphs under the record shape below, and note that a seventh return comes from a different agent afterwards.
They return their record; they do not write it. One writer to the ledger, and it is the context that spawned them, filling each heading in as its answer comes back. N agents appending to one Markdown file lose records to each other, and a ledger holding three of eight claims because two appends collided would grade clean and let the run draft — #206's shared-artifact channel with the sign flipped. Where the harness returns nothing usable, write one file per claim and concatenate; what is not allowed is two writers on one file.
One record per claim, filled in under its heading:
DATE: 2026-08-19
## CLAIM: A white count of 15,000 is within physiologic leukocytosis in pregnancy.
STATUS: sourced
SOURCE: peer-reviewed
REFERENCE: Abbassi-Ghanavati, M., Greer, L. G., & Cunningham, F. G. (2009). Pregnancy and
laboratory studies. Obstetrics and Gynecology, 114(6), 1326-1331.
RESTATEMENT: The table gives a third-trimester white cell range of 5.6 to 16.9 x 10^9/L in
normal pregnancy.
RECENCY: nothing newer - searched 2026-08-19, no later reference-range table for pregnancy exists.
RESOLVED: https://doi.org/10.1097/AOG.0b013e3181c2bde8 - read 2026-08-19
PAGE-YEAR: 2009 - stated on the article's masthead and in the journal citation.
REFUTATION: stands - the volume, issue and pages match the publisher's landing page, and the
third-trimester row is on page 1327.
SECOND-ROUTE: publisher landing page -> journal PDF and table on page 1327
STATED-EXPIRY: none stated
STATUS is sourced or unsourced, and an unsourced record says on the same line what was
searched. SOURCE is one of society guideline, peer-reviewed, government or
tertiary reference. RECENCY is one of current, within five, nothing newer or
guideline in force, and the last two carry the reason after a hyphen — the run must have looked,
and must say so. DATE is the day the paper is written, and the recency rule is measured against
it rather than against the clock. RESOLVED is the URL or DOI the agent actually opened and the
day it opened it — the word read or retrieved, then an ISO date. PAGE-YEAR is the year the
page itself states and where on the page it says so. STATED-EXPIRY is none stated, an ISO date
and where the document states it, or an ISO date followed by superseded cited deliberately and a
reason. Transcribe only what the document states; do not infer an expiry from a publication cadence.
42 C.F.R. § 414.56 (2025) is the known case where none stated is correct: the codification year is
provenance, and the annual reissue schedule is not a stated expiry. REFUTATION is stands,
refuted or paywalled with the reason after a hyphen. A field's value may wrap onto the next line.
The grader also refuses a SECOND-ROUTE with no ASCII -> separator.
It refuses a SECOND-ROUTE with an empty half.
It refuses a SECOND-ROUTE whose normalized halves are equal.
It refuses a STATED-EXPIRY outside the three forms.
It refuses a stated expiry at or before DATE without the deliberate-supersession reason.
Two of those returns are what stops a citation nobody can check, and the third is a second agent. A reference in correct APA form is not evidence that the document exists — an invented one looks like scholarship, which is exactly why a wrong citation is worse than no citation: it survives review.
RESOLVED and PAGE-YEAR come back from the agent that did the research. It was on the page,
so it writes down what it opened, when, and the year the page itself carries along with where the
page says so. No tool here touches the network — the fetching already happened during the
research, and what these two fields do is turn it into something the clinician can audit in one
click instead of a claim nobody can check. PAGE-YEAR has to agree with the year in REFERENCE;
where a source genuinely carries no date, REFERENCE reads n.d. and PAGE-YEAR says the page
states none, and the two agree that way.
Then a refutation pass, by a second agent — not the one that wrote the record. One per sourced
claim, all at once, into the same ledger by the same one writer. The brief is adversarial: here is
a reference and a restatement, try to prove it wrong*.* Not check whether this is right,
because an agent asked that says yes. It looks for the document at the locator, checks the year, the
volume, the numbering and the pages, and reads whether the source says what the restatement says it
says. It also returns SECOND-ROUTE: <research route> -> <refutation route>. The ASCII ->
separator and both substantive halves are required, and the two normalized halves must differ.
Before writing paywalled, try the clinician's authenticated Chrome route through
mcp__claude-in-chrome__*; the in-app Browser pane is not that signed-in route. Refuter
independence remains orchestrator-owned; see research_ledger.DECLARED_LIMITS for the mechanical
boundary. A source is paywalled only when its body remains inaccessible through that
Authenticated route; an anonymous or in-app login wall does not establish the disposition.
Where the profile says the Authenticated route is available, the researcher must try it before
giving up on the intended source, settling for a reachable substitute, or writing
STATUS: unsourced because the body was inaccessible.
It comes back stands, refuted or paywalled, with the reason after a hyphen. A refuted
record is a failure and not an outcome — unlike unsourced, which is honest and goes to
PROPOSED. It means a false citation is sitting in the ledger, so the claim goes back through this
step and comes out either with a sound record or as unsourced. It is never drafted from.
paywalled is the third word, and it exists because a wall is not the same thing as an absence.
A locator that 404s, or that names a document a search cannot find, is refuted — the citation may
be invented, which is the whole failure this pass is for. A live page whose title and authors match
the entry, with the body behind a subscription, is paywalled and passes: the URL resolving to
the right document is itself evidence the document exists, and that is most of what a fabricated
citation cannot do. Say what did match — the title, the authors, the date the page shows.
It is the weakest disposition that passes, and the run says so on its own face. The report
counts paywalled records on their own line, because a set of citations all behind a wall has been
checked far less than a clean exit suggests. It passes because a resolving locator whose title and
authors match the entry is evidence that the document exists, while the separate count preserves
that the source body did not verify the claim. No tool here opens a socket; access belongs to the
research and refutation passes, including the required Authenticated route attempt above.
The independence is an instruction and not a check. Nothing in a record shows which agent wrote it, so the grader cannot tell a real second reading from the first agent answering itself — that is what a written instruction cannot do is fail arriving at its own successor. The one shape the grader does reach is a refutation that is the restatement pasted back.
The ledger is gitignored, because scratch/ is, and that is where a case study's working
material belongs — not in a tracked notes directory. Where the harness ships a general research
skill, borrow the fan-out from it and change that one thing: they write findings into the repo,
and a case study's working material is a patient record.
What makes a record bad, in full, so this can be walked without running anything. A record can be several of them at once:
| The record | Why |
|---|---|
| a field missing or empty | a record missing its restatement is a citation nobody checked |
a STATUS that is neither word |
it decides which of the rules below apply, so a third word is a record graded on nothing |
an unsourced with nothing said about what was searched |
anybody can write unsourced; nobody writes searched PubMed, IDSA and UpToDate without having looked |
an unsourced record carrying a REFERENCE, RESOLVED, PAGE-YEAR or REFUTATION |
the two contradict, and nothing can tell which was meant |
a SOURCE outside the four |
a returned source outside the classes is a finding, not an answer |
a RECENCY outside the four |
it gates the window below, so a fifth word is a record the window never read |
a RESTATEMENT that is the claim pasted back |
the whole point is the source's own terms |
| a claim carrying a number whose restatement carries none | "the source discusses leukocytosis in pregnancy" against a claim about 15,000 cells |
| a reference stating no year | n.d. is legitimate APA and cannot be measured for recency — unless an excuse with a reason stands in for the year |
a reference more than five years before DATE with no excuse |
the amended rule above |
| an excuse with no reason after it | the run must have looked, and must say so |
a RESOLVED that is not a URL or a DOI |
the field exists to put a specific in front of a reader, and on the society website is not one |
a RESOLVED that does not say when it was read |
a topic page changes under its citation, so when matters as much as where |
| a locator read after the paper was written | a record describing a reading that had not happened yet |
a PAGE-YEAR stating no year, against an entry that states one |
the entry claims a year the page did not give |
a PAGE-YEAR that is a year and nothing else |
a year alone is an assertion; where it was found is a place a reader can go and look |
a PAGE-YEAR that is not the year in REFERENCE |
the row a fabricated citation has to get past |
a REFUTATION outside the three |
it gates the row below, so a fourth word is a record the refutation never read |
a REFUTATION with no reason after it |
the run must have looked, and must say so, arriving at the second pass |
a REFUTATION reading refuted |
a false citation is sitting in the ledger: rewrite the record or write unsourced |
a REFUTATION that is the restatement pasted back |
the first agent re-asserting rather than a second one checking |
Two things are deliberately not on that list. Within two years is the target is a target, so a
current disposition on a three-year-old source is not a defect. And an unsourced record is
not a defect at all — it is the honest outcome the PROPOSED block exists for.
Once the prescriptions exist, hand the ledger and draft to a fresh checker as well -- #289's rows read the draft as well as the ledger the way #298's row below reads the evidence dump:
python tools/research_ledger.py <claims-ledger> --draft <the draft>
| The prescription | Why |
|---|---|
| a drug in an Rx table that no claim record names | the dose is the highest-stakes claim in the document and the one every other gate exits 0 on |
| an order stating a dose whose claim record states no number | a record naming the drug is not yet a record that sourced the dose, and this is the form of that a string test reaches |
| a prescription table with no readable drug row | a table this cannot read is a finding and never a table quietly dropped from the set |
Without --draft those rows do not run, and the report prints not graded against them
rather than 0. A zero beside a row that never ran is the silent pass this whole arrangement
exists to refuse, so the run that graded no prescriptions says so on the same page as its clean
exit. A draft carrying no readable prescription table is exit 2 for the same reason.
The choice not to build a dose-correctness table is grounded in indication, weight, renal function,
pregnancy, route, and #215's warning against rejecting a correct result for the wrong reason. The
mechanical boundary of these rows is named only in research_ledger.DECLARED_LIMITS.
Nothing downstream reads the number either, and that stopped being true on
#299. Step 9's the Rx blocks
row asks a reader whether every drug has a table, whether every Sig ends in an indication and
whether the prose block is there; it does not open the ledger and compare the dose, and it still
does not. the dose against the record that sourced it is the row that does — a reader and not a
row here, because a string test can only ask whether the table's number appears in the record, and
1 g against 1000 mg and q24h against once daily are the same unit problem
NUMERIC_CLAIM_UNQUANTIFIED above refuses to touch. Its false-alarm rate could not be grounded
either: when it was ruled, the only run in the tree predated the rows above it and every one of its
prescriptions reached no claim record at all, so there was not one drug-row-and-record pair anywhere
to measure a string test against. #97's
precedent is that a cut point is grounded where the corpus offers one and refused where it does not.
And give a fresh checker the ledger and what you were actually handed -- #298's row, ruled by the clinician 2026-08-20, grades what the run says it read:
First ingest that deliberately supplied file into the shared account-owned store. Use a stable lowercase dump id naming the course, module, and receipt date; never point the ingest command at a directory or let the reporting sweep choose a file:
Ask at intake: Did this dump come with a separate reference list? If yes, retain that exact
supplied file with --references; if no, omit the option. Retention records provenance for the
future primary-source join and does not claim that the join exists today.
python tools/uptodate_store.py ingest <the evidence dump> --dump-id <course-module-date> --module <course and module> --received-on <YYYY-MM-DD> [--references <the supplied reference list>]
The command copies the raw dump and writes its manifest and searchable index under
scratch/uptodate/. The raw source and manifest stay gitignored. Search the accumulated store with
python tools/uptodate_store.py search <query> [<query> ...]; use several literal synonym queries
when needed. python tools/uptodate_store.py sweep only reports topic-shaped material that has not
been deliberately filed and ingests nothing.
python tools/research_ledger.py <claims-ledger> --evidence <the evidence dump>
| The citation | Why |
|---|---|
| an UpToDate topic cited here that no accumulated manifest carries | the deliberately supplied store is the required source set; an unfiled current dump or a topic opened through another route does not put it in that set |
| an UpToDate topic whose literature-review month has left the signed two-year window | re-read it while the profile says the clinician has an account; UPTODATE-ACCOUNT: no waives this row without pretending the old date became current |
| an entry whose locator names an UpToDate topic and that states no database element | the row above reads a topic only from the database element, so without this one an entry missing it escapes the check and the coverage count together |
The grounding is companion-evidence membership, not whether some route can open the page. The Authenticated route may reach an UpToDate topic outside the accumulated manifests; that does not add the topic to the faculty material the clinician supplied. The clinician hands dumps over wholesale, so their manifests accumulate across courses. This accumulated manifest population is the required supplied-source set. A journal article, a society guideline or a government page the dump lacks is left alone, because that is this step's ordinary case: a claim record only exists because the evidence did not cover the claim, and a row firing on those would refuse the correct outcome.
A topic the dump merely refers to and does not carry is not a defect and is not graded. The dump cross-references far more topics than it carries -- by better than an order of magnitude in the one this was measured on -- and the great majority will never be cited. Firing on those would fire on almost every case study, which is the rate at which a warning stops being read.
There is no escape hatch, and that is the ruling rather than an oversight. If an UpToDate topic
is worth citing it goes in the dump, and the remedy for a finding is one paste. The second row is
what keeps that true: the first reads a topic only from the UpToDate. element APA gives it, so an
entry that drops that element was invisible to the check and to the count of what the check read --
four characters, and a citation walks around a row with no hatch. So an entry this cannot read is a
finding, never a citation dropped from the set in silence. Without
--evidence the row does not run and the report prints not graded against it rather than 0,
on the same reasoning as the prescription rows above. An evidence file carrying no topic body at
all is exit 2, because a dump this cannot read would otherwise fire the row on every UpToDate
citation in the ledger -- a mass false finding rather than a scan.
#298 records the declined wider
join and its rationale; the implemented boundary is named only in
research_ledger.DECLARED_LIMITS.
Then hand the ledger to a fresh checker, and do not draft until that checker reports it clean:
python tools/research_ledger.py <claims-ledger>
The grader's coverage boundaries are inventoried in
research_ledger.DECLARED_LIMITS; this skill points there without copying its rows.
Exit 0 is clean, 1 names how many records failed, and 2 means it did not scan — no file, no
records, or no DATE header. Re-run with --show to see which records, and that output is PHI:
read it, do not paste it. The command's full coverage inventory is
research_ledger.DECLARED_LIMITS; the refutation pass remains the clinical-source reading.
Every rule the command applies is written above, so a harness with no Python walks the ledger by
eye instead. The command saves the reading; it is not where the rule lives. That is
icd10-cpt's arrangement with tools/specificity_scan.py, and AGENTS.md keeps
the two classes of tool citation apart deliberately.
Where the harness has no subagent tool, the same briefs are worked one at a time in the main
context, into the same ledger. The mechanism is the ledger and the brief; the parallelism is a
speed property, and the grader cannot tell the difference. This settles
#214's open question 1, and
#218 takes the same answer
rather than inventing a second one. Where the harness cannot research at all — no subagent, no
search, nothing to read — the record is written STATUS: unsourced with that said plainly, and the
deferral behavior is what is left: the claim goes to PROPOSED and, if it is a number, out of the
document. Deferral is the floor when research is impossible, never the choice when it is merely
work.
A claim found unsourced in the middle of drafting goes back through this step, not into the body
with verify this against it. That was the first run's behavior and it is what this step exists to
replace.
4. Write the Sanity Check
Four confirmations, one per line, each ending - confirmed, then Sanity Check completed - proceed. The four are the module or case number, which video, the hyperlink, and a one-line
description of the case.
5. Draft the body
In skeleton order, in his voice — ../_shared/reference/style.md is the authority and the
part that matters most is that the voice is first person and decisive. I would, I will,
I'm going to stop. Never the provider should consider.
Read scratch/voice-model.md first, if it exists, and write each section in the register that
section takes — the MDM, the patient education and the reflective prose are three different voices
and ../_shared/reference/voice.md §2 says which is which. Where the model declares a
register unmodeled, that section is written in the §11 mechanics and the gap is declared in
PROPOSED. Where no model exists at all, this run does not stop to build one — that needs the
clinician and his samples, it is setup-clinical-skills step 8,
and a case study is usually being written against a deadline. Declare it and name the skill.
Two things every MDM entry carries: the discriminator — what in this case puts the diagnosis in or out, not a textbook summary of the disease — and a citation. Ruled-out entries end on the verdict, and the strongest form in the corpus promotes the verdict to the entry's own header line with the reasoning underneath.
6. Write the prescriptions
The fixed six-row table in ../_shared/reference/style.md, one per drug, including the
home medications that are being continued unchanged. The patient cell is a placeholder and the date
of birth is literally x-x-xxx — a case study prescription carries no identifiers. Sig spells
the numbers out and ends for <indication>. Held orders say so in the drug row.
Then a short prose block under each table, carrying the five fields the spec's Pharmacologic Therapy component names and the table does not: drug class, contraindications, monitoring, adverse effects, and the guideline that supports the choice. One paragraph, not a second table. Ruled 2026-08-18. The shape and the worked example are in ../_shared/reference/style.md §8, which is the authority on section shapes and the one place they are written.
Why prose rather than more rows. The spec asks eleven fields per medication and the table carries six, so something had to give. Eleven rows stops it looking like a prescription; nowhere leaves a scored component answered only by accident, in whatever the Patient Education happened to say. The table is where the order belongs and the prose is where graded reasoning belongs, and the guideline citation in that block is the cheapest Integration of Evidence-Based Guidelines point in the document.
Omitting them has never cost a point, which is not the same as being safe — it is the mode finding again, one section down. See Three modes, and none of them subtracts a section above.
Then hand the ledger and draft to a fresh checker, which grades the half of step 3 that could not run before the tables existed:
python tools/research_ledger.py <claims-ledger> --draft <the draft>
Every rule it applies is written out in step 3 above, so a harness with no Python walks the drug rows by eye instead. A drug with no claim record goes back through step 3, not into the document with a citation borrowed from the nearest source that mentions the disease -- #289 is that behavior and it is what this exists to replace.
7. Fix the references
APA 7, alphabetized — ../_shared/reference/apa7.md is the rule, and it is checked
rather than recalled. That sheet carries the a/b disambiguation ordering, the UpToDate entry
form, when a retrieval date belongs and when it is a defect, and the mechanics of the list itself.
It also carries the legal-entry form and points to the code-owned configured-reader boundary.
An APA question it does not answer is looked up at apastyle.apa.org, never guessed.
Roughly alphabetical was a description of the corpus and never the standard. Sorted is
sorted.
This walk is not optional and its findings are not handed back. Ruled 2026-08-18 — see What it is graded by above. Walk the defect list, every time:
| Defect | Fix |
|---|---|
| The reference list head |
Truncated - read the full file at https://github.com/mshamblin5150-code/clinical-skills/blob/ab406eb62505e934c34b036398602e78e885ae25/skills/practicum-case-study/SKILL.md.