Imported from alex-kzr/skills (
software-development/feature-pipeline-operator/SKILL.md). Install upstream withnpx skills add alex-kzr/skills --skill feature-pipeline-operator. Copyright stays with the author.
Feature-pipeline operator
Procedure
- Read the repository pipeline/operator guides, selected task, feature prompt, plan, and predecessor evidence. Before any runner preflight or dry run, compare every selected task's declared
Depends onandSupersedesIDs with the selected plan's declared task graph; correct an unstarted plan/task mismatch through the approved planning workflow before creating run state, because the runner rejects cross-plan dependencies. Establish whether the user selected the universal pipeline or explicitly declined it; an explicit direct-workflow instruction takes precedence over any default pipeline procedure. If the user replaces the requested task before dispatch, discard the earlier selection and recompute the replacement task's dependency closure, reusable evidence, dry run, and preflight; never carry the original task into the run by momentum. Confirm dependency readiness and explicit plan approval only when the pipeline path is selected. - When the user asks to author a plan for a runner behavior change, create aligned named branches in the umbrella repository and its nested core repository before writing. Save the source intent as an English
.prompts/<date>-<slug>.mdfile even though.prompts/is Git-ignored; create the English plan, task files, and active-board cards in the umbrella repository. When a recovery must complete before a stalled source task can resume, place the recovery card first in## To Dowhile preserving the source task's runner-projected active state; stage that one board insertion separately from unrelated projection changes. Parse the plan and execution metadata with the core parser, rungit diff --check, and commit only the plan artifacts in the umbrella repository. Leave the core branch clean until implementation begins. - When the user explicitly declines the universal pipeline, do not launch, inspect through runner commands, or resume it. Preserve any existing runner state and reports unchanged; switch to the repository's direct board workflow in an isolated worktree, record the result as self-validated rather than independently verified, and run every task-declared check before moving the card to Done. If the switch follows a failed pipeline launch that left executor edits in the original checkout, treat those edits as unverified evidence: reproduce and validate the task in the clean isolated worktree, then integrate only the tested task-scoped paths. Commit the nested repository first, update the parent gitlink and direct-board projection second, and leave unrelated dirty paths unstaged. When the user explicitly asks for a narrow runner-code correction without creating or amending a task, make an isolated direct patch: leave all task/plan/board files and
.pipelineartifacts untouched, add a focused regression first, run the affected suite plus declared checks, and commit only the scoped code and parent gitlink. When the user explicitly broadens that task, amend its current task contract and active-board entry before code: name only the added production, regression, and direct-workflow projection paths, then keep all other historical task files and run artifacts untouched. If the task was already completed by a direct workflow, require explicit authorization to reopen it, restore only its active-board andIn Progresspresentation, retain a short direct-workflow amendment note, and replace—not duplicate—its final self-validated## Resultafter the expanded work passes its declared checks. If the user later explicitly requests pipeline verification of a completed direct workflow, preserve the direct self-validation and its card/result unchanged; do not recycle the completed task's lifecycle or hand-edit it back to active. Create an explicitly authorized, verification-only pipeline task with a fresh feature identity and no implementation ownership, then run the standard runner verification through that task so its final PASS/PASS evidence is independent and does not overwrite the direct result. Give that task a zero repair budget, a no-write executor instruction that forbids declared checks (the runner owns them), the source task and historical run artifacts as out of scope, and the full original declared command set; require executor attributionknown-emptyplus fresh PASS/PASS before reporting independent verification. In a direct-workflow Result, identify the isolated validation by commit/evidence references only; never write a host-absolute workspace path into task documentation. For a parent worktree with a local nested Git repository, populate the nested path with a detached worktree at the parent gitlink commit when normal submodule initialization is unavailable; this preserves the same source revision without changing the parent checkout. Before baseline checks or a dry run, verify every task-required ignored input exists in the new worktree:git worktree adddoes not copy ignored.prompts/or a project-local.agents/link. Resolve the authoritative agents root using the repository rule and materialize a directory junction on Windows when symlink privileges are unavailable; restore each missing prompt only from a known authoritative copy and compare its SHA-256 before use. Otherwise, run--mode plan-only --dry-runfirst. For this read-only contract check, pass an explicit adapter if needed but omit--modeland--effort: those controls are execute-mode only. The explicit--dry-runis required to exercise default verified-reuse and supersession resolution; plainplan-onlymay report an unmet declared dependency before that resolution runs. Treat exit10as the expected pending-gate result; inspect selected tasks, execution scope, reused sources, planned dispatch set, adapter, and declared checks. - Preflight the chosen adapter's effective workspace, role resolution, and runtime controls. When authoring a recovery task, validate its proposed literal
Executorbefore adding the card: a taskTypeselects a working root but does not create an adapter role. Reuse a role proven by the selected adapter or declare an inline role that its resolver accepts. Before executing, confirm the selected task's concreteExecutorresolves under the exact routed working root; if an adapter provides roles inline, prove its availability check accepts that inline role rather than assuming an on-disk agent file exists. Confirm it can read the task file, plan, prompt, every required skill, and every declared Allowed-scope path outside its routed working root using the exact generated argv, working directory, directory grants, and role tools that execution will use. For a nested working root, inspect the production-composed executor argv for the minimal cross-root scope directories; task/plan/prompt grants alone do not prove that declared sibling documentation or source paths are reachable. For an adapter that copies or mounts an isolated executor workspace, also prove the complete ownership handoff before a real task run: the launched worker creates a scoped sentinel, then the runner process reads and removes it after the worker exits. Project-root readability and a worker-only write do not establish usable sandbox access; platform sandboxes can apply output ACLs that prevent runner attribution or promotion. When executor or runner-owned verification commands use a package-manager cache, probe that cache under their exact routed working roots before launch; configure the same explicit worktree-local per-run cache environment for both paths rather than changing the user's global cache configuration. A worker-only cache override does not validate the verifier command that later runs outside the worker process. When the assignment specifies a model or reasoning effort, inspect the runner's actual adapter argv/config propagation and establish a per-run control before launch. When the assignment also names a quota-based fallback, first inspect the primary CLI's native quota display and record the current reset windows and consumption; for Claude, open an interactive session and use/usage, then exit without dispatching work. Follow it with one bounded, no-write request using the exact model and effort before opening runner state. Login status alone does not establish capacity, and a usage display alone does not establish the task's isolation, tool, or verifier eligibility. For a Claude CLI that supports these controls, make the probe non-persistent and tool-free with--print --permission-mode plan --no-session-persistence, require an exact short response, and retain its structured success envelope as capacity evidence. A passing bounded probe proves only that one turn is available; it does not establish that the adapter can satisfy the task's isolation, tool, or verifier contract. If an actual primary-adapter pre-launch check rejects strict isolation, tool-free verification, role capability, or another non-quota boundary, preserve that evidence and do not treat the availability probe as authority to switch adapters; obtain an explicit recovery or adapter decision before creating a fallback run. If either the probe or any actual primary-adapter role launch (executor, task verifier, or test verifier) reports a quota or session-limit failure, preserve that evidence and use the requested fallback only in a fresh self-contained feature identity. A verifier quota failure is still adapter-capacity evidence: do not resume the primary identity or misclassify it as a task implementation failure. Before dispatching that fallback, prove its production-composed isolated workspace handoff; if it later fails the complete worker-to-runner ownership handoff, do not weaken its sandbox merely to continue, because selecting the fallback adapter is not authorization for a weaker sandbox. If the user explicitly expands the active task to cover the handoff repair, persist the complete canonical amendment and matching Markdown contract first; then do not redispatch an unchanged runner into the already-proven inaccessible workspace. A task cannot repair a sandbox boundary it cannot enter—retain the amended evidence and obtain an independently executable bootstrap route or adapter choice. If the only demonstrated bootstrap route requires a weaker adapter sandbox, stop for separate explicit user authorization that names the fallback. Limit it to a disposable isolated workspace and write a production-argv regression proving ordinary and read-only launches retain their normal sandbox. For a supported in-task amendment, retain the same task and run revision rather than creating a duplicate feature identity solely because Markdown changed; renew every proof that binds the prior task digest or revision before resume. Re-probe the requested primary with the exact model and effort only when the recovery needs that adapter and capacity is demonstrably restored. If a runner supplies a context bundle instead of filesystem access, inspect the rendered executor prompt and prove it contains the complete digest-bound input contents and explicitly satisfies the task's read requirement; a path listing is not access. For a recovery that changes required-input derivation or launch composition, compile the execution scope before validating or reading any task-required inputs, then derive grants and context only for that scope—never every task parsed from the plan. Cover this order in production composition so an invalid skill declaration on an unselected task cannot prevent an otherwise eligible focused run. For a Claude executor inacceptEditsmode, pre-approve only each shell-free declared check whose CWD is the executor working root as an exactBash(<argv>)grant; leave checks outside that root to runner-owned verification rather than widening the executor workspace. Before dispatch, inspect the actual compiled executor argv and rendered prompt for those exact grants. The rendered executor prompt must not instruct the worker to run, report, or self-verify any declared command that is runner-owned or lacks an exact grant; contradictory prompt prose causes an avoidable sandbox block even when the argv is correctly constrained. If a task repairs that grant composition or blocked-envelope parsing itself and the current runner cannot pass the complete declared check, treat it as a bootstrap failure: do not dispatch the task into a sandbox that cannot validate its own contract; request authorization for a standalone recovery scope with a production launch-composition fixture. Add a production-composition regression with an unselected task whose required skill is intentionally unreadable, then inspect the first real nested-worker launch artifact; direct adapter argv tests do not prove bootstrap wiring. If no supported probe can establish access, stop before execution, preserve the blocker, and present the scoped bootstrap-recovery proposal for explicit user authorization before adding a new plan or board card. If an unstarted recovery task already has the needed production files and tests in its Allowed scope, amend that task's requirements and acceptance criteria after approval instead of creating a second card; only a task without durable run state is safe to amend. Before dispatch, inventory every production lifecycle boundary needed by every acceptance criterion (including the actual owner of verifier evidence, repair-report handoff, prompt construction, and verdict parser—not only state/report modules) and every compatibility fixture the declared suite validates; include all required production and regression paths in one approved task amendment before launch. For any machine-checked text identity, render the production prompt and compare its literal required form with the parser's accepted grammar; semantic instructions such as separate references to two values are insufficient when the parser requires one exact combined token. When a tool-less verifier must judge lifecycle, escalation, or resume criteria, inspect its rendered prompt and byte-identical runner-owned payload before launch; include durable operation history and relevant executor evidence, not only declared-command records, so the verifier can assess every criterion without inventing evidence. When the routed working tree is already dirty and the task declares a full-suite check, run that exact declared suite before creating run state. If it fails because of pre-existing out-of-scope changes, preserve the baseline output and stop for a scoped recovery or clean-worktree decision; otherwise the task burns its repair budget on failures its executor cannot repair. For an acceptance criterion about interruption or resume, trace the production lifecycle entry point and include its lifecycle-level regression file; serializer-only fixtures cannot prove checkpoint recovery or non-duplication. A task rooted in a nested subdirectory can need inputs outside that root. Before delegating to a coding CLI from the nested repository, grant the parent project root explicitly and have the worker read the exact absolute task, run-state, and source paths; a nested-only search can falsely report missing authority and abandon valid work. Treat a coding CLI's claim that it launched background agents as a status report, not completion: inspect the actual diff and run focused plus declared checks yourself. If a required owner is excluded, stop before execute and request a scoped contract amendment; a repair worker must not change an out-of-scope file or manufacture a PASS. - If a run references a missing prompt, inspect prior run records and session evidence for the exact historical prompt path. Restore only a byte-identical copy from an existing source, verify its digest and a dry run, and never synthesize a replacement prompt from a plan or task file.
- Before execute, derive the run feature identity and inspect its existing
.pipeline/runs/<feature>/run.jsonplus task launch artifacts. For a historical resume, pass the exact stored logical--prompt,--plan, and run-directory name with--feature <feature>; do not rely on CLI defaults because the resume identity guard compares those paths byte-for-byte. A plan's default feature can point at a different immutable run with incompatible adapter or runtime controls. After a bootstrap prerequisite passes independent verification, inspect the intended consumer run's pinned adapter, task identity, and recorded bootstrap authority before proposing its next launch. The prerequisite's task-bound proof cannot authorize the consumer, and ordinary resume or a task-contract amendment cannot switch an older run's adapter; name the supported audited recovery or fresh-identity route and obtain separate task-bound authorization before launching. If a reconciliation registry must bind a new run as its own futuresource_run, choose and pass an explicit bare--feature <feature>name before authoring the task, prompt, or mapping; use that exact name in every mapping and confirm the dry-run’s derived identity. Add a pre-dispatch literal check that every registrysource_runequals the selected--featurevalue. Never derivesource_runby adding an adapter or model suffix: registry lookup resolves only the actual run-directory basename. If the default feature directory contains historical verified dependency evidence, choose a distinct fresh feature identity before the primary launch as well as before any quota fallback; otherwise a new nonterminal run can replace the reusable source and make the fallback dependency-ineligible. If a dependency already has implementation evidence or a nonterminal state, stop and preserve it; do not start a fresh full-chain execution over that identity. - When a focused replacement is blocked solely because an upstream task is independently verified inside a terminal
blockedsource run, preserve that source. Implement the narrowly scoped reuse/recovery correction under an explicit human-approved bootstrap scope; never rewrite the source run to look verified. Prove that task-level path, contract digest, both verdicts, and verification time are still required. - For a superseded dependency, first dry-run the normal focused selection and require its execution scope and planned dispatch set to contain the verified replacement rather than the blocked predecessor. Then exercise the production execute selection path with a fake adapter or equivalent composition test; preview-only substitution can leave
executedispatching the retired predecessor. An attestation ofDEP_ID=SOURCE_FEATUREcan only reuse a task record with that exact dependency ID; it cannot substitute a differently named replacement. If either preview or execute selection still selects the blocked predecessor, stop and route a scoped runner correction with a regression test before executing. After that correction, use a new feature identity and an explicit attestation only when its source carries the exact dependency ID. Confirm the dry run resolves the source and schedules only the target before one--mode execute --approve-planprocess. Retain the selected adapter for any resume. - On completion, read
run.json, the executor report, captured command evidence, both verifier reports, and the projected board/task view before reporting a task as verified. Treat the terminalrun.jsonplus final verifier artifacts as authoritative completion evidence. If a task file claimsDoneor PASS/PASS but its referenced run directory or final verifier artifacts are absent, treat the claim as non-reconcilable rather than verified: do not resume its predecessors or hand-edit the board. Before proposing a replacement, search registered Git worktrees and other known local execution copies for the named run directory; inspect the candidaterun.jsonand named final reports, compare the source and consumer task-contract digests, then restore the complete artifact tree byte-for-byte and compare a per-file SHA-256 manifest. Only after that successful restoration may a dry run establish reusable evidence; otherwise obtain a fresh independently verified replacement before runner-owned reconciliation. When triaging stale active cards with a supersession chain, trace each direct replacement to the first claimed terminal source and validate that source’s on-disk run, named final reports, task identity, and verdicts before proposing any registry mapping; a completed Markdown successor does not establish evidence for any ancestor. Treat a runner--statusinspector as untrusted until a production-boundary regression proves it is side-effect free for the installed core: never add execute-mode or approval flags to a live persisted run merely to obtain status. If status evaluates a gate, times out, or changes any run/workspace byte, preserve the evidence, stop using it, and route a separate recovery task; inspectrun.jsondirectly in the meantime. For every selected Markdown-backed task, read the actual board and task file: a terminal completed task must be absent from the active board, have exactly one checkedDonestatus, and have exactly one runner-owned## Result; an in-progress task must be in the active board with the matching status. Treat a mismatch between durable state and these files as a production projection defect even if the runner exits0and both verifier reports say PASS; preserve the run and create a fresh recovery task with a real execute-path projection regression rather than manually changing the board or task file. Invoke the runner from its core directory with absolute filesystem anchors for--project-root,--agents-root, and--core-root; keep profile, plan, prompt, and project-skill arguments logical relative paths. Do not pass--modelor--effortto--status: those controls are execute-only. Compare recorded repair attempts with the task limit and ensure no executor launch or repair artifact exists beyond that limit. If a later whole-worktree check differs from recorded task verification because projection or pre-existing unrelated edits changed the tree, report that separately; recorded runner evidence remains the task verdict, and the operator must not manually rewrite runner-owned projection. - If a run is nonterminal with no live worker or has exhausted its repair budget, preserve every run byte and do not manually rewrite state. Treat a Kanban task as one logical feature: when evidence proves it needs paths or criteria outside its original contract, obtain explicit approval to expand that same task rather than creating a card solely for finishing it. Use the runner's approved in-task amendment lifecycle when available; it must retain prior revisions and verdicts, open a fresh revision-scoped execution/verification epoch, and keep normal resume strict for unapproved drift. Before resuming, inspect the planned verifier artifact identity as well as the task attempt: an amended revision must create fresh immutable verifier prompts, reports, envelopes, and verdict records rather than reuse an earlier revision's
verify-<attempt>artifacts, because their amendment identity is stale even when command evidence is reusable. Before proposing any amendment, inspect both the Markdown task status/result and every matching durable run record. Only a task with no durable run state is safe to amend as unstarted; conflicting Markdown and durable states, or a claimed completion whose cited final reports are absent, are non-reconcilable evidence and require a fresh recovery task rather than an amendment or manual projection. Before amending a live task, compile its current TaskSpec and submit the completecanonical_amendment_fields(spec)payload, not only the field being changed; omitted amendable fields receive different canonical defaults and make strict resume reject the stored amendment. Before resuming, make the canonical Markdown task contract exactly reflect that complete approved payload and prove its digest matches the stored amended revision; a persisted amendment alone cannot make a strict resume accept stale task metadata. If the installed runner lacks that lifecycle, do not fake it by editingrun.json, resetting attempts, resuming an incompatible fingerprint, or launching the same task ID under a new bootstrap feature identity. A user's approval to keep all work in one card authorizes scope, not a duplicate run lifecycle. On authorization, create and prioritize one pipeline-infrastructure task that adds the amendment lifecycle, verify that capability independently, then use it to continue the original task ID. Adapter fallback remains a separate execution identity only when the prior launch failed for quota/session capacity, because adapter pinning is part of a run identity. When implementation must continue in the same worktree, preserve source artifacts and prevent executor writes outside the approved revision scope from reaching the primary worktree; a fresh snapshot must attribute only approved changes. - If an executor reports
blocked, preserve the report and run state. When a Claude executor reports that a declared interpreter or check command was denied by its sandbox, classify it as an adapter grant-composition failure, not a quota failure: do not switch adapters, resume, or rerun the task; retain the executor-owned delta and request authorization for a scoped runner recovery that proves the exact grant reaches the launched worker. A human recovery may transition the task toready, but resume only after confirming the stored plan/task graph is compatible. Do not edit runner state directly or invent dependency evidence. - When adding a standalone recovery plan, make its task graph self-contained: every
Depends onID and everySupersedestarget must be declared in that plan. Model historical evidence as fixture context, not a formal dependency, unless it is eligible under the current canonical task path and contract digest; an ineligible source makes the selected run fail before verification. The runner validates dependencies and supersession targets against the selected plan, not against other plans or historical evidence; record cross-plan blocked evidence as context, not as a formal dependency or supersession declaration. - When the user asks to stop execution and reconcile cards or commit completed work, do not start a new pipeline run. Read the terminal
run.jsonand both final verifier reports for each candidate; use the reports recorded by the task's final accepted verification generation, not an earlier failed repair gate. Move only tasks with two final PASS verdicts to Done, place repair-budget exhaustion with a task-verifier FAIL in Blocked, and leave unfinished tasks active. Run the declared tests andgit diff --check, stage only files attributable to the verified scope, commit code in a nested repository before committing the parent submodule pointer and project documentation, then report unrelated dirty paths separately. If the user requests promotion to a protectedmain, first inspectmain..FEATURE_BRANCHand the merge-base so the merge scope is explicit. For a nested-core change, push and open the core PR first. When the user explicitly chooses the delivery split “core through PR; umbrella through direct push,” update the umbrella gitlink by direct push to the exact core PR-head SHA after scoped validation; do not wait for merge solely to synchronize the parent pointer. Verify both remote refs withgit ls-remotebefore reporting publication. Otherwise, wait until required core checks are complete and successful, then merge before updating the umbrella. After a CLI merge, inspect the nested checkout and parent status: the CLI can leave the nested repository at the mergedmainSHA while the umbrella still records the pre-merge feature SHA. Commit that exact new gitlink on the umbrella feature branch, verify its staged submodule log contains only the merged core change, then open or update the umbrella PR with that pointer. Re-run a transient umbrella promotion check only after the referenced core SHA has completed its checks. Never bypass rulesets. Merge only into a clean main worktree, then checkout the nested repository at the parent-recorded gitlink and verify the main worktree is clean before reporting success.
Pitfalls
- When WSL uses a Windows-mounted HOME to select an existing Codex auth reference, set Docker's CLI configuration explicitly to the WSL installation and verify image inspection under those exact environment variables. Otherwise the inherited Windows Docker context can make a locally present image appear unavailable; do not inspect or copy credential contents to diagnose this.
- A Windows-created Git worktree's
.gitpointer may be unusable by Linux Git even though Windows Git works. Prove Git boundary discovery and read-only status on both the original worktree and the runner's disposable copy, including nested submodules, before treating an executor window as attributable; a successful Codex process withattribution_state=unavailableis not implementation evidence. Do not rewrite source.gitpointers or claim an ad-hoc Git shim as isolation proof. - For a settings-volume verifier's report-to-envelope protocol, normalize CRLF and lone CR to LF before persisting the redacted report, then hash and parse that same normalized text rather than raw model stdout. Test CRLF and synthetic secret-shaped report text on both Windows and Linux: Windows text-mode writes can otherwise add blank lines on readback and make the tool-free envelope's digest disagree with the saved report. Reject oversized reports explicitly instead of truncating them, because a truncated envelope loses its byte/digest binding.
- Before resuming after an interrupted or blocked verifier turn, check whether the attempt/revision directory already contains prompts, reports, or failure records. Do not overwrite those runner-owned artifacts by replaying the same attempt; obtain a lifecycle-supported fresh artifact identity or a separately authorized amendment and renew revision-bound proofs. After each amendment, compare the complete prior canonical contract digest to the durable run, include the new amendment file in Allowed scope, apply with
--mode amend, then renew both executor and per-role verifier proofs before executing. - On WSL, a fresh shell launcher derived from a Windows-written file can have CRLF and lose argument continuations (
\\\r). Verify raw line endings before any credential-bearing launch; thewrite_filehelper may preserve an existing file's newline style even when passed normalized text, so explicitly normalize a scratch launcher with POSIX bytes and verifyCR=0. An apparent shell exit of 0 after broken continuations is not a proof. - Treat the exact runner-declared Linux/WSL suite as the task's gate even if the full Windows suite passes. Linux-only symlink/bind and argv-composition tests may fail or be skipped on Windows. Inspect each failed test as either a real production defect or a stale assertion before changing code; a verifier's semantic FAIL cannot be converted to PASS by a later green Windows run.
- A Docker Claude READY proof does not establish that OAuth will still be valid when semantic verification runs hours later. On a redacted
401 OAuth token expiredreport, preserve historical artifacts and compare approved host credential-file metadata to the Docker settings volume without reading values: a newer/different host file alongside a successful host login can mean a stale contained snapshot, not absent host authorization. Refresh only the previously approved files under existing authorization; read back target metadata, renew both role-specific READY proofs and the revision-bound executor proof, and rerun on unused artifact paths. Equal metadata alone does not prove token validity, and copying unchanged files or replaying an attempt is not recovery. - If both semantic verifier reports and envelopes say PASS but the runner rejects
missing-amendment-justification-finding, keep the formal failure and historical reports. Inspect the redacted report: a heading such as<redacted> finding:lacks the required literal label, often because a prefix likeAuthorizationtriggered secret redaction. Strengthen both verifier prompts to require the standalone, unprefixedAmendment-justification finding:label, add a failing prompt regression before code, and obtain new revision-bound proofs/attempt; never rewrite the report, relax the validator, or count discarded envelopes as PASS/PASS. - For a Docker Claude settings-volume report/envelope protocol, verify session continuity before relying on
--resume: separatedocker run --rmlaunches with a new private tmpfs and--no-session-persistencedo not demonstrate a resumable transcript. If the envelope exits nonzero, preserve its zero-byte/failed artifact and diagnose the actual transport boundary before any new revision or credential-bearing retry; never reuse its attempt directory. - Re-run the core test suite from the nested core root inside its paired umbrella worktree, but run umbrella
git diff --checkand Git-status inspection from the umbrella root. The core suite can resolve parent task fixtures through the paired umbrella, so a standalone core worktree may falsely fail on missing project documents. Do not append a relative nested-repository path after changing into the core, because the anchor resolves to a nonexistent child and can turn a passing verification into a misleading combined-command failure. - When a CI coverage gate differs from a passing local platform run, treat the CI report as the threshold authority: compute the exact deficit from its total and covered units, add deterministic cross-platform tests for portable failure paths, and rerun the complete coverage command before publishing. Do not lower the threshold or rely on a host-only pass, because platform-specific test discovery and branch execution can change the aggregate percentage.
- Treat a plan-only dry run as routing validation, not worker-access or concrete-role-resolution validation, because no executor sandbox or adapter availability check is launched.
- Validate adapter-native
--modeland--effortagainst the production-composed launch argv/config, not merely an--mode execute --dry-runlabel. Inspect C3 first: if it still saysstage implement: plan-onlyand shows no worker argv, the preview proves task selection and gates only; a component-level model flag test or a successful quota probe cannot establish the runner's effective launch controls. - Probe custom executor-role resolution before execute; a generated inline role charter does not prove the adapter's pre-launch resolver accepts that role.
- Stop on a task-input access denial; do not let an executor infer requirements from partial context, because required skills and acceptance criteria are part of the contract.
- Do not create a new recovery plan or Kanban card merely because a selected task exposes a broader runner defect; state the exact blocker and proposed scope, then wait for the user's explicit authorization, because recovery work changes the agreed task graph.
- Do not use an ad-hoc CLI wrapper, an uninspected context bundle, a direct adapter unit test, or an extra directory grant as proof of sandbox access; verify the production-composed, effective launched worker can actually read each required artifact, because component-level argv construction need not be wired into bootstrap or dispatch. Treat a clean probe, a policy-rejected command, or a probe with different write/sandbox controls as diagnostic only; strict capability proof must demonstrate an allowed operation and denial of forbidden access under the exact executor/verifier runtime.
- Do not mutate a user's global CLI model settings to satisfy one run; adapter launches can outlive the change and concurrent work can observe it. Use a runner-supported per-run model/effort control, or stop before dispatch if none exists.
- Before a credential-bearing verifier live probe, load the target durable Run and parse the current selected TaskSpec; require the stored task-contract digest and active revision to match before invoking Docker. A probe requires a persisted selected run and will correctly reject a Markdown/durable-contract mismatch before any container work. If the backend work needs paths or acceptance criteria outside that task, obtain explicit authorization for a complete same-task amendment, or present a separate verification-only task as the alternative; never attach fresh proof to the historical revision or edit
.pipelinemanually. When a Docker live probe persists strict-isolation proof, exercise--resumewith the exact explicit Docker controls used by ordinary invocation. Recompose the adapter from the selected task's durable run record even when no CLI control changed; otherwise the initial fail-closed capabilities survive and the runner never consumes the proof. After an approved amendment changes the selected task's canonical contract, rerun the live probe before resume because proof validation binds the task-contract digest and revision. - Make the authorized live probe the only first operation of a fresh contained bootstrap. Persist task-bound approval, selected execution scope, contract identity, and Docker runtime controls before Docker composition; then allow only a matching resume to hydrate omitted controls and consume that same run’s durable probe. Reject conflicting explicit controls and never let fresh execution dispatch merely because a named approval exists. Parse a model readiness response at the adapter boundary as exactly one JSON object with no duplicate members and a byte-for-byte expected result; reject JSONL, prefixes, trailing data, whitespace variants, and duplicate keys. Expose only a bounded, non-authorizing status on failure, and erase raw stdout, stderr, and session identifiers before returning; diagnostics must not turn model output into runner evidence, logs, or adapter-fallback authority. If a probe exits nonzero with an unknown category, classify the CLI's structured JSON
is_errorand documented error status from itsresultin memory before proposing a different authentication architecture; stderr may be empty while the JSON result contains the decisive API error. Compare a fresh exact-model host probe and credential-source versus provisioned-copy metadata before diagnosing stale auth or capacity. Preserve raw output and credential contents outside logs, chat, and durable evidence. - If a Docker Codex executor exits nonzero, classify only bounded fields from its structured JSONL
error/turn.failedrecords; a usage-limit response is an adapter-capacity failure, not a semantic verdict or repair attempt. Preservelaunch-failure-Nand historical proofs. Before a later retry, re-read the exact durable run, confirm the next executor generation has unused immutable paths, no verifier attempt artifacts or live lease, and revalidate the pinned current-revision proofs; do not amend, repeat a probe, replace auth, or switch adapters merely to retry a quota failure. - For verifier-proof artifact publication that must resist parent replacement, run the runner on a platform with descriptor-relative no-follow operations and prove those primitives plus the parent-replacement regression under the exact runner runtime. If the native host lacks them, use a configured Linux/WSL runner sharing the Docker engine; never substitute lexical path validation, resolve checks, or post-write audits, because a parent can be replaced between those checks and publication.
- Default to a strict verifier-authority boundary: re-read and validate durable proof for the current run, task scope, contract revision, runtime identity, and canonical controls at every direct launch boundary, and never accept a caller-provided capability object, callback, boolean, proof label, artifact path, or synthetic fixture as authority. If the user explicitly chooses runner-process trust instead of a separate OS authority service, make that mode an explicit immutable runtime-control value and report it as a limitation: retain durable proof, fresh preflight, task/run/role binding, replay controls, and container isolation, but never claim resistance to malicious code already executing in the runner process.
- Do not dispatch a normal universal task to build the first verifier backend when the current runner cannot already launch its independent verifier roles. Launchers are composed before executor output is promoted, so the task cannot use its own unverified implementation to settle its own verdicts; obtain explicit direct-bootstrap authorization for the minimal backend or provision and prove the backend first, then use a fresh verification-only pipeline task for independent PASS/PASS.
- Treat verifier live-probe evidence as runner authority only when verifier composition receives the current durable Run, canonical registered artifact reference, and selected execution scope from the runner. Bind the exact run, scope, role, contract revision, runtime fingerprint, and canonical launch controls; never let a proof object choose a run directory, artifact root, or JSON path, because caller-created lookalike state can be forged by an executor or model.
- For concurrent bootstrap recovery, acquire the pipeline lease before opening or saving a resumed run, then re-read and revalidate its recovery authority under that lease; recheck fresh-run identity collisions there too. A contender must not record lease-contention events by saving an unleased stale snapshot, or it may erase a newly committed recovery row and proof.
- Persist task-bound bootstrap authorization in the initial durable controls before opening a dispatchable run. In the core resume lifecycle, derive bootstrap enforcement from the effective adapter as well as controls: reject Docker-adapter/control runtime disagreement and require task-bound recorded provenance for every Docker resume, including callers that bypass CLI composition. Two successful verifier READY probes do not authorize Docker executor resume when an older run lacks
contained_bootstrap; ordinary--mode amendchanges task contracts, not this control. Before proposing in-run recovery, inspect the full executor operation history and distinguish a prior successful worker from a pre-launch failure; any authorization must be prospective and require fresh current-revision executor proof, attribution, commands, and independent verdicts, never retroactive approval. Preserve historical proof artifacts at their referenced digest: a shared live-probe path may already be overwritten by a later revision, so choose and test an immutable revision-scoped path before another probe. Exercise the exact production CLI order through authorization, proof, and resume in a regression; passing state-helper tests can miss an earlier guard that makes the route unreachable. Preserve the run and seek an explicit choice of audited in-run recovery versus a new identity with historical-run reconciliation; never hand-edit state or silently duplicate the task. - Use the process environment (for example
UV_CACHE_DIR) for an ordinary worktree-local package cache. Reserve a runner's--uv-cache-dircontrol for its documented operational-unblock lifecycle; passing it on a normal run can make a later--resumefail closed because the runner treats it as an unblock control. - Do not treat a model or effort named in worker prompt prose as launch configuration; require its exact native argv/config evidence or create a scoped recovery task that adds that runner control.
- Do not re-run a focused task through a fresh feature run unless its verified dependency source is eligible; a terminal blocked run may contribute only task-specific evidence after the reuse policy explicitly validates that task's identity and independent verdicts.
- When reusable-evidence lookup reports an ineligible source but a matching historical source appears to exist, enumerate every candidate's terminal run status, named-task status, canonical task path, contract digest/version, both verdicts, verification time, and durable attempt/revision identity before proposing a repair; lookup diagnostics can report the first denial after no candidate satisfies the complete identity.
- Require named final verifier reports only from records that carry an authoritative attempt/revision identity; legacy records without that identity cannot address a
verify-Nartifact path, so preserve their established PASS/PASS compatibility while rejecting attempt-bearing records whose named final artifacts are absent. - Inspect the derived run identity before
--verify-dependency-chain: an existing dependency launch artifact makes a new full-chain launch collide with immutable evidence, so preserve the state and route runner recovery instead. - Do not use
--verify-dependency-chainmerely to bypass unavailable reuse after a quota fallback; it redispatches the whole closure, can consume upstream repair budgets, and may stop before the requested task ever launches. - Do not fix a bootstrap deadlock by editing
.pipeline/run.jsonor declaring the whole blocked source successful; preserve source bytes and make eligibility depend on the named task's recorded evidence. - Before adding
## Supersessionto an already verified replacement task, recompute its canonical evidence identity: if supersession declarations participate in the contract digest, the edit invalidates the exact PASS/PASS evidence it needs. Add and test a separate validated reconciliation registry or other immutable mapping boundary, then exercise a production reconciliation run against the real evidence store; fixture-only mapping tests do not expose ambiguous or ineligible historical sources. - Do not delete failed-adapter artifacts to force a different adapter onto the same run; quota evidence is operational provenance, while an adapter change requires a verified lifecycle-recovery path with explicit authorization.
- Do not resume after a human recovery if the runner reports a task-set or dependency-graph mismatch; preserve the evidence and route a scoped runner correction instead.
- Inspect the source-run failure class before using
--recover-source; its linked replacement path is adapter-specific and is not a substitute for a standalone, self-contained recovery plan after a terminal executorblockedresult. - Build pre-implementation recovery fixtures from the exact launch-artifact set the runner persists, including its executor diagnostic when a failed launch writes one; an allowlist that omits a runner-owned diagnostic rejects every genuine quota recovery, while a broad allowlist can admit implementation evidence.
- Exercise both a dry run and a fresh focused execute fixture after changing verified-reuse semantics. The dry run may pass a
TaskSpecthrough a different source-identity path than unit fixtures usingTaskDefinition; cover both shapes, and validate amendment-version evidence only when its active revision and immutable amendment digest match the current canonical amendable fields. - Before implementing a human-unblock eligibility rule, construct its regression fixture from the exact persisted
run.json, artifact map, and referenced artifact paths of the blocked run; derive diagnostic locations from the runner-recorded artifact references rather than hard-coding a presumed report layout, because pre-dispatch and verification failures persist different directory shapes. - Do not accept a third repair merely because it might repair the feature: a task above its declared repair budget lacks valid lifecycle evidence and must be treated as blocked until the runner invariant is repaired.
- Before increasing a repair budget, prove the failed verifier requirement has a task-fixable remediation. A verifier must not require a PASS from its own current pass or its companion as input evidence; that is a circular verifier-guidance defect, so repair the prompt/evidence composition with a regression before resuming.
- Treat a board card as one logical feature: do not create a separate completion-only recovery card when verified evidence shows the same task needs more scope. Require explicit approval and an immutable, revision-scoped task amendment; if the installed runner cannot represent that lifecycle yet, preserve the source and prioritize the runner capability rather than manually resetting its fingerprint or repair count. When asking for this lifecycle decision, explain the amendment, preserve-as-blocked, and verification-only alternatives with their evidence and unblock consequences before requesting a choice.
- Distinguish the runner-captured verification snapshot from a post-run global diff: board/task projection and unrelated dirty files can change the latter after task checks ran, so preserve the verified evidence and route any remaining hygiene issue by its own scope.
- Do not mark a card Done from a runner's summary alone; read the final task-verifier and test-verifier reports, because an earlier failed report or a contradictory prose verdict can make a projected status untrustworthy.
- Judge a runner-captured verification command by its recorded argv, working directory, disposition, and exit code, not by isolated
error:text in its stdout or stderr; integration suites may deliberately exercise failing subprocesses while the suite itself exits 0. - Before dispatching a task that adds a regression containing synthetic verification commands, run its focused test once against the clean baseline and ensure every fixture command is a schema-valid shell-free argv; a malformed test fixture fails during
TaskSpecconstruction and can consume the task's bounded repair budget before any production behavior is exercised. - Parse every amended Markdown task card with the core
load_task_specbefore staging it. Keep authorization notes and narrative under normal body headings, not inside the execution-metadata list, because the metadata parser rejects unknown labels and a documentation-only change can otherwise block the recovery before dispatch. - Before enforcing new durable identity markers on a runner-owned artifact, search the focused test tree for every fixture that passes that artifact into the production boundary. Update fixtures that represent valid current evidence with the required headers, and retain separate malformed fixtures that assert the specific rejection; otherwise a correct boundary check breaks unrelated resume and isolated-workspace coverage only during full verification.
- For a registry mapping that binds historical cards to the focused task's own fresh feature identity, execute the focused task first, then resume that exact terminal run with its stored controls before judging reconciliation. The first execution can evaluate reconciliation before its own PASS/PASS evidence exists; the resume re-enters projection without redispatching terminal work. Treat a verifier's pre-projection assessment of a runner-owned projection criterion as pending, not satisfied: acceptance requires the resumed run's reconciliation event plus every named board card and task Result. Read those exact artifacts before committing. When several historical runs contain the same completed task, select a terminal source whose immutable execution scope contains only the stale cards being reconciled; resuming a source that still has another nonterminal task can dispatch unrelated work. A passing fresh-run test can leave historical terminal cards stale when reconciliation never invokes the projection path.
- When runner-owned reconciliation writes or refreshes a completed Markdown Result before the selected task's command gate, add a production-shaped EOF-format regression for both a first render and an existing Result, then exercise the real idempotent projection before dispatch; a trailing blank line in a historical task can survive a formatter-only change and make the root
git diff --checkfail after the repair budget has started. - When reconciliation augments a focused plan with historical task definitions, use that augmented mapping for both supersession-graph construction and evidence lookup, and make absent historical task files a no-op. Otherwise synthetic fixtures crash and the production graph can name a replacement that lookup cannot resolve.
- Before dispatching a task that adds or extends a domain model, trace the imported public type to its defining module and include that module in Allowed scope. The executor cannot widen its contract during implementation; an otherwise passing implementation is blocked before independent verification for any undeclared write.
- When a task tightens a generated-profile or configuration parser while a later task owns generator/config regeneration, require a compatibility fixture for the currently checked-in generated shape and preserve its plan-only parse path until the later migration lands. Otherwise the first executor delta can make the runner unable to resume for its own verification.
- When normal submodule initialization cannot clone the nested repository during a clean-main merge, create a detached nested-core worktree at the parent-recorded gitlink from an existing local core checkout, then verify its HEAD and cleanliness before reporting the merge. This preserves the merged source revision without inventing or changing remote submodule configuration.
- When a direct-workflow task cannot import or run in a clean worktree because its codebase already contains an uncommitted predecessor delta, characterize the missing contract first. Transfer that delta only after explicit user authorization, amend the current task's scope to name the minimum production paths and its regression/golden coverage, and retain historical run artifacts unchanged; otherwise the task absorbs unreviewed unrelated work.
- Do not sweep all dirty files into a completion commit; stage by verified task scope and commit a nested repository before its parent pointer, because unrelated work must remain independently reviewable. Before staging the parent gitlink, compare its recorded core SHA with the proposed nested commit using
git merge-base --is-ancestorand inspect symmetric history. If the branch diverged, do not silently repoint the parent to the new core commit or use anoursmerge to hide predecessor changes; leave parent docs/gitlink uncommitted until the divergent commits are reconciled and validated, while preserving unrelated dirty files. - When executor work runs in a copied isolated workspace, exclude every runner-generated artifact root from both the opening snapshot and closing attribution boundary. The runner may mirror reports only after the executor exits, and an unexcluded mirror is falsely charged as an executor out-of-scope write.
- For isolated-workspace promotion, test all three provenance cases before dispatch: preserve runner-owned projections and unrelated user-dirty/untracked paths byte-for-byte, but promote an allowed pre-existing path only when the active task owns its prior attribution and that attribution's recorded primary-baseline digest matches the current primary bytes. Verify the primary-worktree digest matches the attributed workspace result before transitioning to
implemented; a blanket skip discards valid task output, while a blanket promotion overwrites unrelated work. - When the direct board workflow validates implementation already present in its starting revision, record the outcome as self-validated evidence without inventing implementation attribution; stage only the board/task Result projection, because the source change belongs to its earlier provenance.
- When a one-shot scheduled pipeline run is overdue or its scheduler listing still says
scheduled, inspect the target run'srun.jsonand live worker processes before firing it manually; scheduler metadata can lag a started tick, and a duplicate dispatch can create conflicting run evidence. - When transplanting a task stash to a clean branch for a fresh feature identity, retain the approved contract and source deltas but restore any stale runner-projected task or board state to
To Do; a new run must create its ownIn Progress,Done, and Result projection rather than inheriting lifecycle claims from an incompatible run. Keep untracked amendment artifacts as historical context unless the new task contract explicitly includes them; they cannot establish the new run's revision identity.
