Imported from arianjad/skills (
thinker-worker/SKILL.md). Install upstream withnpx skills add arianjad/skills --skill thinker-worker. Copyright stays with the author.
Thinker/worker
Use this mode for the current task when the user invokes it. Preserve a manually chosen front-end model and effort. If that choice conflicts with the intended thinker role, explain the conflict; only an explicit user instruction deactivates the mode. Do not change global defaults or infer that this skill switched the running model.
The coordinator (Astra in Codex, Fable in Claude Code) owns framing, planning, dispatch, adversarial review, changed assumptions, and final acceptance. Sol, Opus, or Sonnet owns substantial bounded execution. The coordinator may read sources, run deterministic checks, and handle small bookkeeping directly. Reuse a compliant worker where useful; before continuation, check its actual identity, role, model/effort evidence, and assignment. An unknown or preactivation child is not qualified by a new prompt. Workers do not launch children unless the coordinator has explicitly authorized a bounded exception; the fresh-dispatch guard does not prove caller identity on every surface.
Keep briefs compact: task, owned files, input sources, constraints, check and stopping condition, checkpoint requirement, and whether the worker can edit. A worker returns exact artifacts, checks, and uncertainty. The coordinator spot-checks load-bearing claims and accepts or redirects. Escalate a changed objective, scientific assumption, or consequential finding; ordinary fixes stay with the worker. Independent review is optional unless requested or justified by risk, and requires a bounded authorization cited in its brief.
Fresh worker briefs start with exactly TW-Role: worker. A narrowly checkable optional Codex Luna or Claude Sonnet task starts with TW-Role: leaf; an authorized independent review starts with TW-Role: independent-review; authorized ideation starts with TW-Role: ideation. Review and ideation are separate dispatches with separate briefs: review is convergent (what is wrong, with evidence); ideation is divergent (at most N ranked directions, each labeled speculative with what would confirm or kill it). Both carry TW-Authorization: and TW-Scope: lines in brief lines 2–12. A review brief carries the artifact at a fixed commit or path, the acceptance criteria (a recompute check such as "recompute X from the raw data" is a criterion), and the scope, and none of the coordinator's own reasoning: no expected answers, no explanation of the artifact, no question that points at a suspected gap. The reviewer's first pass is an open "what is wrong". Coordinator hypotheses, if any, go to a second fresh dispatch after that report or to a worker verification task, never into the review brief (evidence: arXiv 2603.12123, one study, moderate effect; re-evaluate against our own review receipts). Lines 2–12 of every plaintext brief also carry a routing header, each key once: TW-Class: (an effortmining class: T1-mechanical, T2-simple-transform, T3-moderate-reasoning, T4-hard-reasoning, R-research, C-coding), TW-Deliverable: and TW-Accept: (one line each), and TW-Risk: (none, or distinct values from destructive, external, physics); each value at most 1000 characters (VALUE_MAX in tw.py). The guard denies a brief without it. Allowed models and tiers per role come from routes.json; the tier is the child's effort. The default model per role is routes.json router.defaults, printed by activate as default models: … (tw.py models --harness <h> marks each role's default); use it unless the task or budget argues otherwise. Which models each role admits, at which tiers (a role's model_tiers entry narrows its tiers for one model, in the guard and the router's ladder), with which priors, which is the default, and how each is dispatched: run python "${CLAUDE_SKILL_DIR}/scripts/tw.py" models --harness claude (one JSON line per role and model; this list is not repeated in these docs, so it cannot go stale). A via: native child is a tw-<role>-<tier> dispatch with no model argument when it is the role's first model (the agent file's pin), else with the per-call model alias; a via: tw.py codex child runs through the Codex pipeline below, and the guard denies a native Agent dispatch that names one. To add, update, or archive a model, follow Adding a model. Pick the tier from the priors activate prints for the brief's TW-Class; deviate by judgment when the brief is clearly easier or harder than its class; physics or convention judgment never below medium; high or xhigh for adversarial reasoning (the router's risk floor lifts only its own pick, never yours). The priors live in one place, routes.json router.priors (a class without an entry takes *; the result is clamped into the role's tiers), merged key by key with the user override file ~/.thinker-worker/routes.json (or TW_ROUTES_OVERRIDE), which may set only router.priors, router.model_priors, router.classes, router.risk_floor, router.defaults and the decision-model keys (router.backends, router.combine, router.options, per-backend kind: "jev" blocks) and survives reinstalls (change a default there, e.g. {"router": {"defaults": {"independent-review": "fable"}}}); a bad override is ignored, logged as an error row (where: "override") on every dispatch, and named by activate/status. Each route row records prior_tier; in Claude Code a dispatch the router takes no action on, at a tier other than its prior, gets one additionalContext line (thinker-worker: prior for <class> is <tier>; you dispatched <tier> (fine if deliberate)). After admission a router logs its own tier pick; the mode per TW-Class comes from routes.json router.classes: shadow only logs, advisory denies with a different tier, active rewrites the dispatch, and a TW-Override: <reason> line in lines 2–12 keeps yours. Decision models decide and you are the fallback: every backend in router.backends (a routes.json block such as kev or laya, kind: "jev", a local POST /v1/systemone server; enable one with the user override, e.g. {"router": {"backends": ["kev"], "classes": {"*": {"mode": "active"}}}}, which lists the backend and runs every class active with its explore kept: decision models rewrite a dispatch only in active (in advisory a gated pick only denies with advice, in shadow it is only logged), and the competing-writer guard can still hold a Claude class advisory) is asked in parallel within router.budget_s, their probabilities are combined by a weighted geometric mean, and the combined pick acts only if its top-1 minus top-2 probability clears router.combine.margin; otherwise your tier stands. Every model's probabilities are logged on the route row either way (details in the harness reference). An explicit user request for a model or effort always wins: add TW-Pin: user <what was asked> in lines 2–12, and the router never advises, rewrites, or explores that dispatch (native or Codex pipeline), and it gets no prior reminder; its route row has pinned: true and promote leaves it out. As shipped the router has no backend and every class is advisory with explore: {"c": 0.5, "power": 0.25, "floor": 0.05}, a decaying schedule: at a class's t-th routed dispatch (t = 1 + its earlier non-pinned route rows across all receipt files) the exploration probability is ε_t = max(floor, min(1, c / t^power)), so 0.5 at first, 0.25 at t = 16, never below 0.05 (SLARouter's forced exploration, arXiv 2606.19376; a plain number is a constant ε): a dispatch above the role's cheapest tier whose ticket (a hash of the brief, without its TW-Route: and TW-Override: lines) falls below ε_t is advised one tier down; every other dispatch runs at your tier. With a backend configured, exploration goes one tier below the final pick, the models' or yours. The router's pick is raised to the highest router.risk_floor among the brief's TW-Risk flags (as shipped, physics, destructive and external all → medium); a pick raised to your tier or above takes no action. In active, the rewrite reaches the coordinator only as the hook's additionalContext line (thinker-worker: dispatched as <router agent> instead of <requested agent> …); the child's real agent is its meta.json agentType, and the route row has action: "rewrite" and router_agent. After accepting or rejecting a child's result, label it: tw.py outcome --harness <h> --session <id> --tool-use-id <id> --accepted yes|no, adding --cause tier|brief|other on a rejection (the tool-use id is on the background-agent notification, or in the subagent's meta.json as toolUseId). Whenever acceptance is mechanically checkable, add TW-Check: <shell command> in lines 2–12 (optional, once, at most 1000 characters); the command must exercise the delivered artifact, not only the child's own new tests. After the child returns, run tw.py outcome: it runs the check once in the dispatch's working directory under bash --noprofile --norc -eo pipefail (Git for Windows' own bash on Windows, never a PATH bash or cmd.exe) and records pass, fail, or unknown, which outranks your --accepted (a weak label, then optional; --no-check skips the check, --check-timeout defaults to 900 s). The check runs in the dispatch's recorded working directory, not the repository you mean, so write it to cd there by absolute path first, e.g. TW-Check: cd /c/Users/Arian/Code/skills && python thinker-worker/scripts/test_tw_check.py. When the child must write one file, name it with TW-Output: <path> in lines 2–12 (optional, once). Claude Code refuses a native subagent's Write to a report/summary/findings/analysis*.md basename, so the guard denies a native dispatch whose TW-Output has one (the tw.py codex pipeline is exempt); name deliverables for their task, e.g. worker-notes.md. Native agents are told to return the complete artifact verbatim if a Write is refused anyway. outcome refuses (exit 2, nothing recorded, no check run) a dispatch that never ran: denied by the guard, advised by the router, or a Codex pipeline run with no cost row. These are prompt declarations, not native tool fields and not proof of task semantics. Do not claim model identity from a worker's self-report.
Exploration protocol.
- A denial saying
router picks <tier> (exploration)means: re-dispatch the same brief unchanged at that agent. The same ticket reuses the cached decision, so it is admitted. - Judge the result as usual. Rejected because the tier was too low →
tw.py outcome … --accepted no --cause tier, then re-dispatch the same brief unchanged at your original tier (admitted: cached, no re-exploration). Rejected for another reason →--cause briefor--cause other. - If the brief must change before a re-dispatch (at the explored tier or back at yours), add
TW-Override: exploration re-dispatch, brief changedin lines 2–12. A changed brief is a new ticket, so withoutTW-Overridea re-dispatch at t−1 can itself be explored down to t−2, and one at t can be explored again. - Keep your tier with
TW-Override: <reason>only where exploration is unsafe for a reason the risk flags do not capture; an overridden dispatch the router would have explored shows on its route row aseligible: truewithaction: null(receipts never hold the override text).
Read Codex routing when using Codex, or Claude routing when using Claude Code. Both describe session activation, native dispatch fields, and what runtime evidence must be checked. A guard admits only fresh native dispatch in an activated exact session. Accepted calls remain subject to normal permissions. Installation, activation, hook trust/loading, live interception, and effective child model/effort are separate observations.
For Codex MultiAgentV2, the native PreToolUse payload encrypts the brief. The guard can check visible model/effort and fork mode, but cannot verify the role line or review scope. The coordinator must check the actual brief; see Codex routing for the reduced guard contract.
In Claude Code, this skill's rendered content supplies the exact current session ID and installed skill directory. To request activation in that session, run python "${CLAUDE_SKILL_DIR}/scripts/tw.py" activate --harness claude --session "${CLAUDE_SESSION_ID}". It takes no role flags; add --store-bodies only to opt into keeping each brief body under ~/.thinker-worker/bodies/ (receipts never hold it). These substitutions apply here in SKILL.md; do not expect them to expand in a separately read reference.
Codex pipeline (Claude Code coordinator). Write the brief to a file, then run, as a background Bash call, python "${CLAUDE_SKILL_DIR}/scripts/tw.py" codex --session "${CLAUDE_SESSION_ID}" --role <role> --tier <tier> --model <model> --brief-file <brief> --cd <repo root>. There is no output-path option: the child writes its deliverable where the brief's TW-Deliverable says. The model must be in both the Claude role's models and some Codex role's models in routes.json (the via: tw.py codex lines of tw.py models --harness claude). It applies the same gate as a native dispatch (activated session, role line, TW-Authorization/TW-Scope for review and ideation, routing header, role tiers and models) and writes a dispatch receipt; a denial exits 2 without calling Codex. Admitted, it runs codex --search exec -m <model> -c model_reasoning_effort=<tier> -s workspace-write in --cd with the role's instructions followed by the brief on stdin, so the child can read, edit, run checks, write temp files, and search the web, like a native child. It prints one JSON line (tool_use_id, last_message, exit_code, effective_model, effective_effort, thread_id) and writes a cost receipt row whose model/effort come from the Codex rollout's turn_context, which is evidence of the model that ran. last_message is only Codex's closing message (codex exec -o, always ~/.thinker-worker/codex/<tool_use_id>.last.md, so it can never overwrite the deliverable); read the deliverable itself, then label it with tw.py outcome --harness claude --session "${CLAUDE_SESSION_ID}" --tool-use-id codex-… --accepted yes|no (a TW-Check runs in --cd). Dispatch Codex models only through this command; a raw codex exec bypasses the gate and the receipts. After admission the router acts as on a native dispatch, with the same route row (plus via: "codex-exec") and the same TW-Override: rule: in shadow the child runs at your tier; in advisory, when the router advises, the command exits 2 with the reason naming --tier <target> and does not call Codex (re-run with that tier, or keep yours with TW-Override: <reason>); in active the child runs at the router's tier and one line before the JSON says so (thinker-worker: dispatched as --tier <target> instead of --tier <yours> …). There is no competing-writer guard on this path; after an active rewrite the command writes a race row with lost true unless the rollout's effective effort equals the router's tier (unknown effort counts as lost), so only a verified rewrite joins promote's router arm. The activate command writes activation state and prints the effective tier priors and risk floors (Tier priors (<source>): …; status prints the same line), so check hook loading and child metadata separately.
