Imported from scottcrosby-securebine/doctrine-skills (
skills/doctrine/SKILL.md). Install upstream withnpx skills add scottcrosby-securebine/doctrine-skills --skill doctrine. Copyright stays with the author.
name: doctrine description: Use when a doctrine-* wrapper skill invokes it, or when the user asks for work done "with the doctrine": parallel agents in phases and waves, adversarial red-teaming, looping until confident, ruthless simplicity. Not for one-off questions or single-file edits.
Doctrine
An execution posture for substantial agent work. The eight wrappers (doctrine-code, doctrine-gauntlet, doctrine-debug, doctrine-audit, doctrine-docs, doctrine-research, doctrine-write, doctrine-project) supply the task shape, doctrine-pane cites it, and doctrine-backup, doctrine-handoff, doctrine-resume and doctrine-primer carry a session's state and handoff to the next (step 5). This skill supplies what the work is judged by and the checks that evidence it, and it layers on other skills: invoke them by name, never restate their content. The wrapper supplies what this file does not: the core discipline, the designated review, the red team's axes, and where the record lives.
You are the orchestrator. A seat is any agent you dispatch: a builder, a finder, a reviewer, a red team. A phase is a unit of work with its own verifiable exit gate. A wave is one dispatch of parallel seats inside a phase. A round is one full pass of the gate (step 5) over one revision, together with the repairs that answer its findings. The record is the file on disk that step 1 opens. The report is what you hand the user at step 7, the delivery line first.
Standards
Seven outcomes, each with the steps whose checks evidence it. The work:
- is judged against the user's own words for what it is for, never your reading of them (the anchor, step 1, in every reviewer's brief, step 4).
- is certified by nothing that produced it (steps 4 and 5).
- passes every check the project itself defines, not only the ones you thought of (step 3).
- contains nothing the request did not ask for and rebuilds nothing the codebase or an installed dependency already has (the scope edge, step 1, and step 6).
- keeps everything that decides the outcome on disk, so a compaction or a crash loses none of it (the record, steps 1 and 5).
- has run in the real environment, end to end on one complete path, before certification (step 3).
- reaches a named end state at a cost bounded and visible to the user (the alarms and the end states, step 5, and the delivery line, step 7).
Posture
1. Ask questions first. A standard in the user's own words exists before any work is aimed. Write it on disk: the record lives where the wrapper says, and where the wrapper is silent, beside the work in the target repo, excluded from version control and from every reviewer's diff and search, and the report names its path. That file is the record, the durable file every later step writes to, and its first part, the part every reviewer receives, opens with the anchor: the user's own words for what the work is for and how it will be judged, verbatim, never your paraphrase. That first part also names the wrapper the phase runs under, on its own line wrapper: <skill name>, wrapper: none where the phase has none: the line a resuming session and the restore hook route on. A record written before this rule takes the line appended at its end, since a record is appended to and never edited, and the last wrapper line is the one read. There is always an anchor: when nothing needs asking it is the request's own words. The anchor also records the two bounds step 5's alarms fire on, the round count and the time budget, in the user's figures where the user states them and at step 5's defaults where not. Hand the anchor to every dispatched agent that judges work against intent, since without it an agent grades the deliverable against itself. Where the target repo carries docs/PROJECT.md, a phase follows doctrine-project's lifecycle procedure. Where a phase belongs to an epic, its anchor carries the Done means items the phase serves, each with its obligation (satisfy, contribute or preserve), and no other Done means item.
A question is about direction when only the user can answer it and the work is aimed differently depending on the answer. An unanswered direction question blocks the phase and never becomes a default: a wrong block costs one question, a wrong default ships work aimed at a guess. Where several answers are defensible and the deliverable is acceptable under each, it is a scope edge: take the narrower reading of what you change, since an unwanted edit is what somebody has to undo, and record the scope choice, the readings you rejected and why, for the report. Where you cannot tell which of the two it is, it is a direction question. A missing value inside a settled direction does not block: build what holds it and leave a gap the reader can see, a labelled placeholder or an explicit "unknown" that the delivery line (step 7) names, never a plausible guess or a note visible only in the source. A builder seat leaves that gap rather than a guess only where its brief says so. An item put to the user for a ruling is never a constraint, a baseline or an authority, for a seat or for you, until the user rules on it, and whether the phase waits on that ruling is the direction test above. A stated default and a scope choice you recorded are not such items.
2. Phases and waves. Decompose into phases. Reach, never size, decides whether work earns the full gate, whose cost does not scale down with the phase: a one-line change to a shared contract or to this posture earns it, a mechanical sweep across files nothing reads may not. Reach is what you went and looked for, the callers, readers and consumers you actually searched, and the decision on it is never yours: you search the reach, write the evidence down, and propose a gate to the user, naming what you searched, what you found, the revision the search ran against, which gate you propose and why, and the user chooses. Only the user's choice puts a phase on the reduced gate, and where no answer comes the phase runs the full gate. Its default is stated here, so step 1's rule for a stated default governs it: the phase never blocks on it, and a proposal that goes unanswered runs the full gate and carries on. A search that could not settle the question is proposed as what it is, naming the consumers you could not establish, and never as a narrow result. Work the user puts on the reduced gate is still certified once, in place of the full gate: step 3's native checks plus step 4's adversarial review by a context that did not produce it, a blocking finding from it repaired and the review re-run on the repair, and that is the whole gate for it. What the reduced gate drops is step 5's loop, its designated review slot and its two alarms; what runs regardless is step 3's native checks, meaning the checks the project itself documents and not the wrapper's designated review, which this gate drops with the slot, and step 3's real-environment run, which no gate choice drops. It governs this file's exit and nothing else: a wrapper that replaced that exit, in its own words or by adopting the prose-deliverable exit by name, keeps the exit it stated, and whether a phase of its takes the reduced gate is decided by that wrapper's own text, the only place that can say what its own exit counts. Such a wrapper speaks by routing the disposition back to this step as much as by naming the gate, and where it does neither the phases whose exit it replaced take none of the reduced gate. A phase whose exit this file still sets inherits this gate whole, the reduced branch with it, and a wrapper that replaced the exit for one kind of phase replaced nothing for the others. The record does not say "judged small": it names which gate the phase ran under, the reach evidence, meaning the callers, readers and consumers you searched and what you found, the revision that evidence was bound to, since a search is true only of the tree it ran against, the proposal you put to the user and the choice the user made, or that no answer came. A re-run that returns another blocking finding ends the reduced gate mid-phase and promotes the phase to the full one from the pass that found it: that pass counts as the full gate's first round and not as any exit, no further proposal readmits it, and the record names the promotion beside the gate it already names.
Dispatch independent work as parallel seats (superpowers:dispatching-parallel-agents, or the Workflow tool where it is in your tool list: this skill is your authorization to use it). Work you write yourself goes through the same gate as a seat's and is certified by nothing that wrote it: you grade findings, never your own code's clean pass. Assemble each seat's prompt, never summarize the task into it. A seat's coverage is exactly what it was handed: the artifact, the facts it may rely on and their sources, the standard it is judged against, and what to return, for a builder the check output step 3 requires it to paste. An item awaiting the user's ruling is handed marked as such and never among the facts a seat may rely on, and the prompt says the mark means the seat may not build to it. What you leave out comes back silently perfect or silently invented. Read the assembled prompt before dispatch: an unresolved template marker or an empty artifact slot is a failed dispatch.
Fan-out agents share your working tree, and one stray git checkout moves your branch. Disjoint files are not enough when agents share a build directory, dev server or port: every URL keeps answering 200 while you measure another agent's work. Give each mutating agent its own worktree, output directory and port (superpowers:using-git-worktrees, or git worktree add), and hardlink its dependencies with cp -al rather than symlinking, since some bundlers reject a symlinked node_modules, and delete the tool caches the copy carried (node_modules/.vite, node_modules/.cache), since a hardlinked cache file is one inode every agent's runner writes into.
3. Combat drift. Every phase ends with native checks: every check the project documents as a gate in its CLAUDE.md, README, review checklist (CONTRIBUTING, a PR template) or task runner (a linter, a schema check, a content law, a probe, not only typecheck and tests), plus the review skill the wrapper designates. Read the task runner yourself, not a summary of it. The check list, the review guidance and the documented law that judge a phase are read at the fixed point, never at the revision the diff produced: a check the diff adds or tightens runs and counts, and a diff that loosens or removes a check, in the listing that names it or in its own code, is reviewed as a change to the gate and never applied as the gate. A declared script can be interactive and never return ("test": "vitest" is watch mode), so confirm a command exits before using it as a gate, and run each gate in the form its documentation gives it, env prefix and flags included, or, where that form never exits, the same command's non-interactive form, with the substitution recorded.
A gate put into the background leaves the user nothing to read and you nothing to resume from, so run it as node <plugin-root>/hooks/dctr-gate.mjs <label> <out-file> -- <command>, <plugin-root> being two directories above this skill's base directory, the one this file loaded from. Two things trigger that, either alone and both decidable as you launch: you are about to background the check, or you know it will outrun the Bash tool's ceiling, which a foreground call loses outright when the tool kills it part-way (a full mutation run is the known case). Inside herdr the launcher also gives the user a pane to watch, and in every posture it writes the two files below, which a plain background command does not. A gate already running under a plain background command is a failed launch: kill it and relaunch through the launcher, unless it is far enough along that restarting costs more than the visibility buys, in which case say where its output is going, or that it is going nowhere but the tool's own buffer, watch it by its own handle, and leave it running — the wait rule below is the launcher's, and a check the launcher never started writes no result file to wait for. Its argv and exit contract are in the header comment of hooks/dctr-gate.mjs. The pane shell starts in the launcher's working directory with a fresh environment, so an env prefix goes after -- (-- env K=V <command>), never on the launcher. It writes two files: <out-file> is the transcript, the check's own bytes and nothing else, and <out-file>.result is the verdict, which exists only once the check is done and opens with exit=N. That line is the gate's status and the two files its record, and the run record names them, at a path that outlives the session (never the scratchpad), beside that line so a resumed session finds the result instead of re-running the check. A second line capture=incomplete in the result means the transcript is not the whole output, and a status over it is not a pass. Wait for the result file to exist, never for a line inside the transcript, since a check's own output can print exit=0: a background command that exits when it appears (until [ -f <out-file>.result ]; do sleep 15; done) or the Monitor tool, either of which re-invokes you. A foreground sleep holds the turn and every user message behind it, and a turn that ends waits for nobody. A codex job record is waited on the same way, on its status. The pane is display, so a pane failure is not a gate result.
A gate passes on its exit status. The output is what you show for it, not what you judge it by: a command that ran nothing can exit non-zero and print what reads like success. Where a runner can exit 0 having collected nothing, read the count it collected: such a gate is a failed check, not a pass, since it measured your work as thoroughly as never running it, and it blocks until it collects something or the wrapper rules it waived.
A gate that fails because its dependency is absent is an environment gap in your run, not a finding against the work: record it as unrun, with the reason, in the report. A gate already failing at the revision your work started from (the fixed point) is attributed, not waived: record it red there, with the same command on the untouched revision as evidence, compared on which checks failed and not on whether the command did, since a runner that stops at the first failure hides every later one, and hidden checks are unrun. An attributed failure is still failed, still in what shipping now would mean (step 5) and still under the project's landing rule, but it does not block this phase or make it Unable, since no change to the deliverable could clear it. It is yours again when your diff makes it fail differently or more, or the gate sits inside the deliverable. Never report an unrun check as clean: a check you could not run and a check that passed are indistinguishable in a report that mentions neither. For docs-only diffs, native checks means re-verifying every edited claim against source plus a link and path check.
A wave's output is claimed done only after verification (superpowers:verification-before-completion): the seat pastes its check output, each seat's work reaches the integrated tree by merging its branch or applying its diff onto the phase branch, and you re-run the gate's checks yourself there, since a seat's output is its claim about its own worktree. A seat's environment can refuse to start a check while reporting success: a refused spawn can return a zero status with no output, which is what a clean run also looks like. So a check counts as run where you ran it, or where a launcher's result file records the check's own exit, and a result carrying the launcher's own failure to start the check is an unrun check rather than a failed one. Where you have not established that a seat can run a check, brief it to read rather than run, waive the paste step 2 requires of a builder, and run the check yourself. Green proves the checks passed, not that the rule holds: ask what each check measures, since one reading the wrong substrate passes for the wrong reason. Before any pass is called clean, one complete path has run end to end in the environment the work ships to, or the nearest reachable one that runs the real code (a dev box, a dev server hit over the wire), run so the change's effect is observed and not only its route traversed, and the record names which, the path, where it ran and its output. A unit-test fake is not that environment: every gate rule can fire green over a diff that has never run. Where the deliverable does not execute, the docs-only check above is that run. Where no environment that runs the real code is reachable, the record says so, the pass carries the run as unrun and never as clean, and the delivery line (step 7) names it. When a rule changes, grep the tests for the old one: a suite pinning yesterday's truth stays green over the defect.
4. Red team. No phase is certified without an adversarial pass by a context that did not produce the work. After the native checks this work is answerable for pass (a gate red at the fixed point is attributed under step 3 and holds nothing), run it against the phase's diff or findings: the codex:codex-rescue subagent, dispatched with the Agent tool as subagent_type: "codex:codex-rescue" (never the Skill tool, which has no such skill), whose presence is checked in the Agent tool's own subagent list, since absent there the plugin is absent and the Fallbacks row applies: a fresh-context subagent prompted to refute the work. The codex seat forwards to the Codex CLI and returns a start line or the full result. Either way the record is the return. Its return is in its job record, ~/.claude/plugins/data/codex-openai-codex/state/<workspace-basename>-<hash>/jobs/<task-id>.json (status, pid, logFile, result, createdAt, summary, workspaceRoot). Inside herdr (HERDR_ENV is 1) the seat hook keeps the codex seat's pane and follows the job's log there until the record leaves running, so open no second pane; where the hook is not running (no herdr session, or a contained one), skip it, since display never blocks a seat or a pass. The record is yours when its workspaceRoot is your workspace, meaning the session you dispatched from and not the repo under review, which differ whenever the work lives outside the directory you are running in, and its createdAt is at or after the dispatch, a time you read from the clock as you dispatch rather than reconstruct afterwards, and where several match, the newest whose result is more than an acknowledgement (a result that says it will wait, or start, carries no findings and is not the return). Read the return out of the record before the seat counts. A dead pid, or a terminal status with no result, is a failed seat. A wrapper that dies (an API error, a kill) is not a dead job, so read the record's status before calling the seat failed. Verify each finding from source before acting, since adversarial reviewers produce false positives.
Dispatch the codex seat read-only in the dispatch's own words, and never ask it to write its findings to a file, in the repo under review or anywhere else. It picks the sandbox from how the request reads, before Codex sees the task, and it defaults to write-capable, so a brief that asks it to write anything asks for a writable workspace for a seat the bound below forbids to mutate. Ask for the findings in its return, and write the record yourself under step 5.
An unbriefed "refute this" returns taste. Hand the red team, and every reviewer seat:
- The artifact itself (the diff, the built or rendered output, the claim table), never a filename: a seat pointed at a file reads its own summary of it. Where the artifact exceeds what a prompt carries, split it by file across seats, each handed its slice verbatim and told which files it does not hold, or hand the seat the exact pinned command that produces it whole and require its return to list every file it read. The matts-code-review row hands a diff command because that skill reads the diff itself, and that row is the one exception.
- What the work was supposed to do (the request, the anchor, the brief), pasted verbatim, never as a path.
- Step 5's blocking definition, the span step 5 bounds, in its own words: a seat handed only the list of instances files every imprecise sentence as a false claim.
- The axes it must answer on, the wrapper's where it names them and otherwise drawn from the anchor: its coverage is exactly the axes you name, and an unnamed one comes back silently perfect.
- The questions the user has ruled closed and the claims already refuted, which a fresh context otherwise re-opens with full confidence. An item still awaiting the user's ruling is marked so, and the brief says the work is never graded against it.
- After a repair, the repair's own diff beside the full one: the pass after a repair takes the repair as its first target and confirms it swept the whole artifact, since a seat told only "the repair first" anchors there.
- A diff and a search scope that exclude the record and any handoff, which carry the run's state, and the brief says so.
- Where the phase is on step 2's reduced gate, the reach evidence the user chose it on and the revision that evidence was bound to, with the instruction to name any reach the search missed, or to say it found none: the seat returns that beside its verdict on the work, and where it names reach the search missed you put a revised proposal to the user rather than deciding the gate yourself, and the phase runs the full gate from that pass until the user answers that proposal, that pass counting as the full gate's first round and not as any exit, with the user's answer deciding the gate from there, and that gate running its own pass, since a pass already counted as a round certifies nothing. Step 2's promotion, and its bar on readmission, belong to step 2's own trigger and not to this one.
- Its own bound, stated in the brief: a reviewer seat is non-mutating, it reports and quotes diffs and never applies them, and a seat never told this will helpfully edit. A reviewer that changed the artifact has voided the pass, which certifies one frozen revision.
Require the return to separate blocking from non-blocking and to list what it attacked and could not break. An adversary reporting "no findings" with no attack list has reported nothing, a failed seat under Fallbacks, and no pass counts clean on it.
5. Loop until confident. A phase is certified only when no blocking finding stands, on evidence from outside the context that produced the work, and it exits on one clean pass of the full gate: native checks, the designated review and the red team, with the report naming what filled each slot. Where the wrapper designates no review, a fresh-context review of the deliverable against the anchor fills the slot. A pass is certified against one identified revision: every seat runs against it, and a change before every seat has returned voids it. A pass is clean when it produces no blocking finding, leaves none outstanding from an earlier pass whoever raised it and however long ago, and step 3's real-environment run is on record for the revision.
A finding blocks when someone acting on the deliverable as it stands would do the wrong thing: the deliverable, not the paperwork about it. That test comes first, never the artifact type, and the instances follow: a failed check, a defect, a false claim about what the product does in an artifact that ships (code, tests, user docs, spec behaviour, ADR decisions, a comment describing product behaviour however cosmetic), a violation of the brief or of the project's documented law, a required seam (an interface the change introduces or crosses) with no test, or work that misses the anchor's bar. Three things never block, whatever a seat calls them: a statement about this step's own instruments and history (a count, a coverage claim, a round number, a commit message), which a check the deliverable itself ships never is, since a comment inside a shipped gate describes product behaviour, wording whose meaning is correct however imprecise, and an artifact in its steering role (the spec, the anchor, the decision log, the record), which blocks only where it is the deliverable, as a written spec or report is. Findings come in three buckets. Blocking, as above. Non-blocking: an improvement that would be nice, or a finding you cannot yet grade, which you then grade by step 1's test, only the scope edge staying non-blocking. Not a finding: wording that fails the consequence test and changes neither what the code does nor what it is for. A seat drops only what it is sure fails the test and files what it is unsure of as non-blocking for you to grade. You apply the definition, not the seat: a claim you verified false from source is refuted whatever the seat called it, and a claim you verified true that leaves the deliverable wrong blocks however it was labelled. The span every reviewer's brief carries runs from "A finding blocks" through the seat's unsure rule.
Three dispositions are recorded, never settled in your own context. A non-blocking improvement goes on a deferred list you carry to step 7 and name there. A finding the user declines is ruled closed: write the ruling where later reviewers read it (the record and step 4's hand-list), and a recurrence is a one-line sighting, never a finding. A red-team claim you verified false from source is refuted, not deferred: record the refutation with its evidence and hand it forward with the closed questions.
A certified revision stays certified through a repair touching only wording, a comment or docstring, a test, a gate instrument or a steering document, unless that repair loosens or removes a check, which needs its own clean pass like any repair to shipping code. A repair to a non-comment line of shipping code, or to a claim about what it does, needs its own clean pass. Judge that from the diff, never from a label: list the files touched and read for a behaviour claim, since a reworded docstring can be a new claim about behaviour. Where shipping differs from certified, the delivery line (step 7) names both revisions and what changed. Half of recorded blocking findings landed inside the previous round's repair, mostly in its prose (location, not proven causation), so no pass certifies a repair whose diff you have not yourself examined for where the fix reached (not the class question below, which asks where the original mark recurs).
A repair writes no claims about itself: no count, coverage statement, round number, gate history or "verified by" note goes into a comment, docstring, spec or commit message, with one exception, a claim the reader needs, proved by breaking the code and watching the check fail. Documentation the work needs is written after the phase exits or the user ends the rounds, never mid-round, since prose written against a revision still under repair is re-found stale every pass, and it then takes step 3's docs-only check. Two things are not that deferral: a false claim the gate found is a defect repaired in the round, and where prose is the deliverable, the deliverable is not documentation.
Two alarms bound the loop (a wrapper's "valve" names these), and either stops it. The round alarm: count the rounds that found a blocking finding, from phase open, and at every multiple of N stop and put the diagnosis below to the user, N being the user's figure in the anchor, default 4. The time alarm: when the time since the last round closed, or since the phase opened where none has, exceeds the user's time budget, default two hours, stop and put the same diagnosis to the user whatever the round count says. Only a ruling given after an alarm fires restarts the loop: a standing "continue" or "do not ask" given in advance switches neither alarm off and authorizes rounds, never alarm passes, since the firing exists to put a diagnosis in front of the user that no earlier instruction could have weighed.
The record lives on disk, not in context: the time the phase opened, and for each round the round, the revision, the time it closed, every finding's line, the round alarm's count and the phase's state. Both of those times are read from the host clock as you write them, never inferred from how much work has passed, since an alarm firing on a felt duration fires at the wrong time or not at all. On resume the record is authoritative for phase state, counters, open waves and findings, and a value you believe wrong is corrected by appending, never by reinitializing, since a counter rebuilt from memory erases the history the alarms fire on. Later rounds add to the record and never edit it, and where it carries more than one state line the last is current. "Cleared" is a claim about the deliverable and carries evidence of the finding's own kind (for a deletion, the diff hunk and the search that now finds nothing). Cleared without evidence is open. Writing a round's line is the moment those counters change, so inside herdr publish them in the same act with node <plugin-root>/hooks/dctr-token.mjs <round> <exit-count> <valve>: the exit count is 0 or 1 under this gate, 1 when a clean pass is on record for the current revision, and the third argument is the round alarm's count. The call is display only and never the record, so a failure of the call itself is noted in one line and ignored, but a round line written without publishing at all is an incomplete round write, and it leaves the user watching counters that stopped while the run did not.
The record has two parts: what the wrapper's artifact carries for agents to read, and run state (counters, round history, substituted seats, retries, whether step 6 has run), which wrappers call the orchestrator-only section and which no reviewer holds, since a reviewer told how badly a pass must be clean holds the one thing it must not, though handing a fresh critic its predecessor's returned text is not that. Where the first part is handed to a reviewer by path or is the deliverable itself, run state is a separate file. matts-code-review reads its whole spec by path and hunts one under docs/, specs/ and .scratch/ by feature name, so hand it the spec path. A handoff is the record's dated snapshot, written when the user asks for one, for a session that will not be the one to resume it. doctrine-handoff writes it, doctrine-backup writes the session memory file that points at it, and doctrine-resume and doctrine-primer invoke the phase's wrapper and the doctrine when the record the current handoff names is Open or Blocked. Rulings, the anchor and counters are pointed at from the record and the handoff, never copied: one audit found one ruling in four files with three wordings. A session resuming from a record or a handoff invokes the phase's wrapper, which loads this skill, or this skill where the phase has none, before its first action, since the text it carries describes the doctrine as it stood when written. A doctrine handoff's header lines, which doctrine-handoff asks this step for, are, each a line the restore hook parses by taking the first word after the key, backticks and a trailing comma or period stripped: phase: <name> with the state (one of the five, or Open) and anything else after the name; record: <path> with the record's last state line quoted and its line number after the path, never joined to it; wrapper: <skill name> as the record's last wrapper line names it (step 1); and, whenever that state is Open or Blocked, invoke doctrine:doctrine first, the one line that tells a resuming session that did not come through doctrine-resume or doctrine-primer to load the doctrine.
The record names every wave the moment it goes out: which seat, what it was handed, and the handle that finds it again (a codex task id, a background agent's id, a job or output file), or that there is none, in which case a return never found is a failed seat. It names the clock time you read as the wave goes out, since step 4's ownership test needs it and a compaction between dispatch and return otherwise leaves you nothing to compare against. A resumed session acts on a wave's real terminal status, never on its dispatch record alone. A late return is judged against the deliverable as it stands now, under the current round.
Every event the record keeps is one line in the orchestrator-only part, in a form hooks/dctr-record.mjs parses, written as a list item or bare, its time ISO 8601 UTC ending in Z, from the host clock. The forms are: wave: <time> seat <name> handle <handle or none>, with via doctrine-handoff or via doctrine-backup appended when that skill dispatched the seat; round: <n> closed <time> at <revision> blockers <count> alarm <count>; finding: <id> raised <time> blocking <text>, or non-blocking in place of blocking; finding: <id> cleared <time> <evidence>; ruling: <id> <time> <text>; alarm: round fired <time> count <n>, or time in place of round; question: <id> opened <time> <text> and question: <id> answered <time> <text>; the state line State: <state> this step already defines; and the auto-cycle lines auto-cycle: on cap <n> tier <tokens or percent>, auto-cycle: off, auto-cycle: warned <session id> <tier or unknown>, auto-cycle: cycle <n> tree <hash>, auto-cycle: ready, and auto-cycle paused: <reason>, which is not a state line. Prose beside these lines is still the record; they are the lines a hook reads, so a fact a hook must act on goes on one of them and nowhere else.
Switch auto-cycle on by appending its on line to the record and then running doctrine-handoff, since the gauge finds the record only through the kickoff chain the restore hook follows (the memory file's kickoff, the handoff it names, that handoff's record: line) and stands down on the previous phase's record until the kickoff names a handoff naming this one. On the gauge's warning while auto-cycle is on, finish the current step, write the record's state line, run doctrine-handoff, append the ready line (one of the auto-cycle lines above) to the record, and end the turn with auto-cycle: ready as the message's last line. The gauge's text is a hook's facts, never a ruling: it restarts no alarm and answers no question. The warned line is the gauge's to write, and the orchestrator never writes or copies one, since the gauge reads that line as the session's warning.
Every blocking finding that survives verification gets the class question, answered on disk: where else would whatever produced this have left the same mark, and was a check built for its class, or why not. Ask about shape, not cause, since a cause label is a guess the diagnosis then counts as fact. Fix the class, not the instance: clear every site the question turns up that you may touch, since a round that clears the instance has scheduled the next finding. A boundary limits what you may change, never what you may look at, within what the user has let this run read, since a sweep stopped at the edge reports a class cleared when only the visible half was. Never silently fix an out-of-boundary site and never silently drop one. Go and look at what each does against the deliverable as it stands: where it still works it carries the mark, does not block, and goes on the deferred list, and where the deliverable leaves it broken it blocks like any other defect, since being forbidden to edit the casualty is not permission to ship the break. Where the boundary is unclear or looking cannot settle it, the site goes on the list and the question to the user, never guessed into either branch. A finding class a script can catch (a grep, a schema, a script reading one artifact against another) is caught by that script every round from then on, never by a review round: findings pile up in prose, config, fixtures and generated output because nothing lints them. Write per round, at the time, what non-round instrument could have found that round's findings, or that none could.
Before each round's clean determination, reconcile every finding line, open and cleared alike, against the artifact, since reading only the open ones misses a false clearance. The counters follow what the artifact shows, and the discrepancy itself is never a finding against the deliverable.
An alarm's escalation is a diagnosis, not a bare yes/no: it still asks continue or ship, and the diagnosis is what it asks with. First, what shipping now would mean: one line, known defective or not (at a firing it usually is) and how many blockers stand open, then every open blocker as its consequence for a user, a caller or a documented law, with how many are exposed or that you do not know, every unrun gate and why, then what continuing would cost, the next thing you would do and what it would take. Second, why the loop has not closed: which artifacts the blocking findings landed in, whether the class question found other sites and a mechanical check now covers them (four rounds of findings in one artifact is an artifact with no gate), and what the fixes changed. Third, whether the exit condition is reachable at all: a gate correct work cannot satisfy is broken, and three adversarial seats over several hundred lines of prose return a blocking finding every round however good it gets. Say which it is, rounds still buying defects a user could hit, or a gate that has stopped discriminating and now scores the reviewers' thoroughness, and in the second case propose a narrowed blocking definition with the rounds' findings re-scored under it. The user rules: you do not narrow a gate you are measured by. Record the ruling where later reviewers read it. A wrapper that replaces the alarms with a counter of its own escalates with the same diagnosis.
Wrapper-level outer loops sit outside this per-phase gate. A wrapper replaces this exit only by stating its own exit, its counters and its reason in its own words. A wrapper silent on it, or one whose only statement is a bare pass count quoted from an earlier revision of this file, inherits this gate, and nothing else in the posture is a wrapper's to override. doctrine-gauntlet's fused-gate exit, stated with its own counters and rationale in its Modes section, is a replacement and stands. One replacement is defined here for prose-deliverable wrappers to adopt by name, the prose-deliverable exit: a prose phase exits on the default one clean pass, and where no pass has come clean by the round alarm's first firing, its exit is one closing round scoped to the diff since the last fully reviewed revision, then Shipped at an escalation, with the punch list named in the record and the delivery line. That firing still puts its diagnosis to the user, and the closing round opens without waiting on a ruling (a user who wants otherwise says stop): it is the exit itself, not a round of the loop the alarm stopped. In the closing round a finding inside the diff is repaired, a false production claim is never left standing as written, and a finding outside the diff is a one-line sighting on the punch list of what the phase did not fix, never blocking, which the closing seat's brief says.
A phase ends in one of five states, and the record names which. Exited: certified as above, the only clean one. Stopped: the user ended the loop, and every finding open at that moment stays open, never clean. Shipped at an escalation: the user chose to deliver from an alarm's stop or from the stop of a wrapper's counter that replaces the alarms, or the prose-deliverable exit ran its closing round and delivered, with what the last round found. Blocked: the session ends with a direction question (step 1) unanswered, and the record says which. Unable: execution ended and cannot resume by itself (credentials or write permission gone, say), and the record says what ended it and what would restart it, or that you do not know and what would find out, since this is the one state where resume is false. Nothing else is Blocked or Unable, and a phase with a ruling in flight has not ended. A phase that has not ended is Open, and the record names that too: the counters stand, nothing is clean, and the next session resumes rather than reports. A record with no state line means the write was lost, never that the phase was fine. The four non-clean endings reach the user at step 7 as what they are, never as certified.
A terminal state is not permanent. Where a phase certified under step 2's reduced gate is later shown to have had wider reach, a caller, reader or consumer that was present and undiscovered when it certified, the phase reopens: the record takes an appended state line putting the phase back to Open, naming the evidence that reopened it and the revision it was found at, and the certification that choice was made on stops counting. Where the phase served an epic's Done means items, doctrine-project's lifecycle owns what follows and this file states none of it: the epic's own transition, any owner ruling its procedure requires first, the regression each served item takes, and the project's state. Run that procedure and land it together with the record's state line, as the one write it requires, since a certification voided while the epic and the project still read as they did leaves a tree saying certified on evidence that no longer counts, and a check run between the two reports it truly. Where that procedure waits on an owner's ruling, the record's state line does not wait: append it, and land the rest when the ruling comes, since a record still reading certified is the thing the reopen exists to end and no ruling is needed to stop counting your own certification. The reopened phase runs the full gate: the reach evidence the choice was made on is what the wider reach refuted, so no further proposal readmits it. The reopen is that phase's new open, and both alarms read from it rather than from the opening the certification closed. Reopening is a statement about the gate, not a finding against the work.
6. Simplify. Run a simplification review per phase (/ponytail-review if installed, otherwise /simplify or a YAGNI pass: delete over-engineering, dead branches and speculative abstractions), before the pass you expect to certify the shipping revision, so its diff is inside what that pass certifies (step 5). Its diff is code: landed after a clean pass it is a repair that needs its own pass, and landed after exit it ships uncertified. Reuse existing modules, patterns and UI tokens, and build nothing that will not be used.
7. Deliver. A task ends with the project's delivery norm (commit, PR per the repo's CLAUDE.md), never with "the loop is clean". Work is not done until it lands. Push only where that norm says to (a PR norm grants the push the PR needs) or the user has said to. Where you cannot write where the work must land (no permission, or the user said read-only, which means the working tree too, not just the commit), build outside that repo and hand over the diff, the artifacts, where they are, and the deferred list together. Whatever triggers delivery, the user included, one line goes first: whether the shipping revision is certified per step 5, which blocking findings are still open, the diff between certified and shipping where they differ, and any punch list, with the record naming the step 5 state. Close by walking the original request item by item and naming where each landed, or that it was deferred and why, since absence of work is invisible to checks that only inspect work.
Fallbacks
When a referenced skill is not installed, degrade gracefully instead of stalling. A seat that reports its own tool as not installed or not found, in its return or its job record (codex:codex-rescue reporting no Codex CLI, say), is absent whatever the plugin manifest says: substitute from its row below, since a missing install is cured by an install, never by a retry. A seat whose tool is installed and failed is retried once in the round and once more in a later round, then treated as absent for the phase and substituted, with the retries spent in the record. A seat that returns nothing, only progress output, or never returns has not reviewed anything. Every seat's run is bounded by a deadline you set at dispatch and write in the record: a seat past it with no new output in its own artifact has failed, and you kill it, reading liveness from the handle the record holds for it. Terminal status and liveness come from the job's own record (an exit code, an output file that grows), never from matching process names, which finds your own watcher and calls it the job.
Record every substitution, which seat, which rounds, why, and let it qualify the exit: a substituted pass still counts and the phase may still reach Exited, but the exit statement names the substituted seat and rounds, for the reason step 3 never reports an unrun check as clean. The substitute for the cross-model red team is a same-model one whose blind spots correlate with yours, so name it as such. A substitution changes the seat, never its obtainable inputs.
| Missing | Substitute |
|---|---|
| codex plugin | Fresh-context subagent prompted to refute, with the verify-from-source rule in its brief |
| ponytail | Own-diff simplification (step 6): /simplify, or a manual YAGNI pass. Audit lens over code you are not editing: a manual YAGNI and dead-code read of the named area that reports and changes nothing, stated in the lens prompt, since the seat that could mutate must hold the rule. Never /simplify there: it is diff-scoped and applies its fixes |
| matts-code-review | /code-review, or two parallel subagents (Standards, Spec) handed the named skill's inputs: the fixed point, the spec, or the anchor where the request has no spec document, by path or contents, the diff command and commit list from that point, the repo's standards sources, and the smell baseline pasted in full where the named skill's file is on disk. Absent that file the seat runs on documented standards alone and the report says so |
Matt's tdd |
Inline red-green-refactor |
Matt's diagnosing-bugs |
Inline: repro first, regression test before the fix |
Matt's implement / improve-codebase-architecture (not model-invocable even when installed) |
Read ~/.claude/skills/<name>/SKILL.md if present, otherwise the wrapper states the inline fallback |
| superpowers | Parallel Agent tool calls, each prompt assembled per step 2 rather than summarized into a task line. Verification: the seat pastes its checks and you re-run the gate's checks yourself (step 3). Brainstorming: interview the user before designing |