Imported from korchasa/flowai (
framework/AGENTS.md). Install upstream withnpx skills add korchasa/flowai --skill framework. Copyright stays with the author.
Framework (Product)
Source of truth for end-user packs (skills, agents) distributed via flowai.
Responsibility
<pack>/pack.yaml— Pack manifest (name, version, description, scaffolds).<pack>/commands/— User-only workflows (SKILL.mddirectories). Names:flowai-*,flowai-setup-*. Source files MUST NOT declaredisable-model-invocation; the CLI writer injects it at sync time based on directory placement. Benchmark scenarios co-located in<pack>/commands/<command>/acceptance-tests/.<pack>/skills/— Agent-invocable capabilities (SKILL.mddirectories). Names:flowai-*. Source files MUST NOT declaredisable-model-invocation. Benchmark scenarios co-located in<pack>/skills/<skill>/acceptance-tests/. Each skill carries TWO categories of scenarios:- Execution scenarios (1+ per skill): verify the skill produces correct results when triggered.
- Trigger scenarios (FR-ACCEPT.TRIGGER, exactly 3 per skill): verify description-matching correctness. One positive (
trigger-pos-1/mod.ts) the skill should activate on; one adjacent-negative (trigger-adj-1/mod.ts) where a different skill is the right match; one false-use-negative (trigger-false-1/mod.ts) inside the skill's domain but with the wrong intent. Coverage gated byscripts/check-trigger-coverage.ts. Authoring guidance inwrite-agent-benchmarks§6.1.
<pack>/agents/— Canonical agent definitions (<agent-name>.mdfiles, IDE-agnostic). Each agent hasname+descriptionfrontmatter and a shared system prompt body. IDE-specific transformation is handled by flowai at install time. Benchmark scenarios co-located in<pack>/agents/<agent-name>/acceptance-tests/.<pack>/assets/— Shared templates (AGENTS.md templates) used by multiple skills and the acceptance test runner.<pack>/acceptance-tests/— Pack-level acceptance test scenarios (e.g., AGENTS.md rules verification) with shared fixtures.
Installation shape
Both <pack>/commands/ and <pack>/skills/ install into the same target directory .{ide}/skills/. The distinction is the disable-model-invocation: true flag on commands (injected by the writer, not authored). IDE-level native slash-command directories (.{ide}/commands/) are reserved for user-owned primitives managed by flowai user-sync — framework commands do not land there.
Packs
core— Base commands (commit, plan, review, init, etc.) + core agents.devtools— Skill/agent authoring tools.engineering— Procedural engineering knowledge (deep-research, write-prd, etc.).deno— Deno-specific skills.typescript— TypeScript-specific setup skills.memex— Memex: long-term knowledge bank for AI agents (three skills: save, ask, audit).beta— Opt-in beta capabilities not yet promoted to core. Ships thedoc-anchors-validateStop hook (Claude-only), theselect-llm-modelskill, and the cross-IDE delegation primitives (ai-ide-runner,delegate-to-ideskills +workeragent, consolidated here from the former standaloneide-bridgepack). With skills present the pack is no longer hook-only — it emits a Codex manifest/marketplace entry too (hasSkillstrue).
Key Decisions
- Scripts in
<pack>/skills/*/scripts/and<pack>/commands/*/scripts/must be standalone-runnable: a Deno script usesjsr:specifiers and no import maps; a Python script uses the standard library only. Prefer Python for a script a skill step tells the user's agent to run — the framework may not assume Deno is installed there, and a script that reaches for the network to make up for it has already cost a measured run (draw-mermaid-diagrams, 2026-08-31: the old validator shelled out tonpx @mermaid-js/mermaid-cli, the cold download blew a 120 s timeout, and the agent shipped a broken diagram). - Nothing shipped under
<pack>/skills/<name>/or<pack>/commands/<name>/may spell out this repo's own documentation layout:check-skills.ts(FR-UNIVERSAL.DOC-SCHEMA) rejects the literalsdocuments/tasks/,documents/requirements.mdanddocuments/design.mdin every file there, including bundledscripts/*.pyand their*_test.ts— onlyacceptance-tests/paths are exempt. Write "thetasksrole from AGENTS.md" in prose, and in a test build the path from parts (["documents", "tasks"].join("/")) when a fixture needs the default. Observed 2026-09-04 ontasks-overview: SKILL.md and the unit test both tripped it on the firstdeno task check. - Skills and commands follow agentskills.io standard; the
commands/vsskills/directory is the framework-level classifier for user-only vs agent-invocable intent. - Agent format is canonical (IDE-agnostic); flowai adds IDE-specific frontmatter during distribution.
- Scaffolded artifact mapping declared in
pack.yamlscaffolds:field (primitive-name → artifact paths; resolves the same whether the primitive lives undercommands/orskills/).
IDE Behavior Notes
- Claude Code skill routing has two surfaces: the model's preloaded skill catalog (frontmatter
description) AND a runtime directory listing of.claude/skills/that the agent can read mid-conversation. The latter lets the agent invoke aSkilltool call by directory name even when the description does not match the query. Implications: (1) skill descriptions cannot fully suppress activation; the directory name is also a routing key. (2) Trigger acceptance tests (FR-ACCEPT.TRIGGER) must treat "agent readsSKILL.mdto answer a meta-question about the skill" as a positive, not a false-use trap. (3) Description quality controls preference among installed skills, not absolute availability.
Composite Skill Authoring (FR-SKILL-COMPOSE)
Composite and atomic SKILL.md files are gitignored build artefacts — never hand-edit, and don't expect to see them in git. Source of truth: framework/atoms/<name>.md atom files + framework/composites/<name>.md composite wrappers + the framework/composites.yaml manifest, all consumed by scripts/generate-skill-composites.ts. Every consumer regenerates first: deno task check, deno task acceptance-tests, deno task build-plugins, and the CI Build framework tarball step each run --write as a prerequisite before reading SKILL.md, so the rendered output is always current and there is no tracked rendered copy that can drift. Manual regeneration: deno run -A scripts/generate-skill-composites.ts --write (idempotent). The 9 generated paths are explicitly listed in .gitignore; the generator's checkGitignoreParity fails the build if the list goes out of sync with --list-targets.
The rules below describe what the generator emits and what its canon validator enforces. They are written for the maintainer who edits framework/atoms/*.md or framework/composites/*.md, not for the agent that runs the SKILL.md.
- No delegation (machine-enforced): Generated SKILL.md MUST contain a
**No delegation**rule in<rules>instructing the agent to execute inlined steps directly and NOT invoke the source primitives via the Skill tool. The wrapper's<rules>block supplies this text; the canon validator rejects the composite if it is missing. - No source-skill names in description (machine-enforced): Frontmatter
description:MUST NOT name any atom from the manifest. The canon validator greps the description againstframework/composites.yamlatoms:keys and fails the build on any match. Naming source skills in the description prompts the model classifier to invoke them via Skill tool. - Self-contained marker (machine-enforced): Description MUST include the exact phrase "Self-contained — execute the inlined steps directly" so the agent reliably treats the composite as terminal. Canon validator rejects composites that drop this phrase.
- Verdict Gate explicitness (machine-enforced): Inter-phase verdict gates MUST tell the agent what to do on EACH verdict, including the success case (e.g., "Approve → DO NOT commit yet. Continue with the Commit steps below."). Gate text lives in
<gate after="N">…</gate>blocks inside the wrapper; the canon validator requires Approve AND a reject path (Request Changes / Needs Discussion / Reject) in the same gate. - One
<step_by_step>per atom slot: Each atom source MUST contain exactly one<step_by_step>block; the atom canon validator rejects multiple. The composite renderer extracts that single block per phase and emits it under a### <title>heading. - Token budget: rendered composite SKILL.md must stay under
SKILL_MAX_LINES(currently 700), atom source must stay underATOM_MAX_LINES(currently 1000). Single source of truth:scripts/lib/skill-limits.ts— both the generator canon andscripts/check-skills.tsimport from there; bump in ONE file, not three. The 10000-token cap fromFR-UNIVERSAL.DISCLOSURE(SKILL_MAX_TOKENS) is relaxed for composites — the list of exempt skills is derived live fromframework/composites.yamlviascripts/lib/composite-list.ts. Adding a composite to the manifest automatically exempts it; there is no separate list to keep in sync. - Size budget — for atom editors: rendered composite SKILL.md size = wrapper base + Σ (each atom's
<step_by_step>block size) + per-phase heading. Edits to an atom file OUTSIDE its<step_by_step>block (frontmatter, Context, Rules, Verification) do NOT inflate composites. When growing an atom, estimate impact:awk '/<step_by_step>/,/<\/step_by_step>/' framework/atoms/<id>.md | wc -l× (number of composites consuming the atom) + currentwc -l framework/core/{skills,commands}/<composite>/SKILL.md. HittingSKILL_MAX_LINESin any composite is the moment to consider compressing the step_by_step block versus bumping the cap. - Atom phrasing — anchor pronouns explicitly (authoring rule): LLM-driven execution resolves pronouns like "SAME / above / earlier / this" against the most recent context entry, which may differ from the human-intended antecedent. When writing or editing atoms:
- "Run the SAME command" → "Run the SAME
test/checkcommand from step 2a" - "as above" → "as in step N"
- "this directory" → "the session-id'd scratch dir from Rule 11"
- "the same way" → name the prior step explicitly.
Verify via acceptance tests: every step that says "SAME / above / earlier" should have at least one scenario covering execution under that step. The
review-no-change-no-alarmscenario caught exactly this ambiguity in the JiT-subset Step 2b (the agent re-interpreted "SAME command" as the JiT-tests, not the project test command).
- "Run the SAME command" → "Run the SAME
- No wrapper-level params (machine-enforced):
_params:blocks are allowed only insideframework/atoms/*.md. Declaring_params:in aframework/composites/*.mdwrapper frontmatter makes the generator fail withwrappers MUST NOT declare _params: (use atom-level params instead)(seescripts/generate-skill-composites.tsfunctionrenderCompositeTarget). Architectural consequence: composites CANNOT be parametrized at the wrapper level (e.g., "ship with-or-without Plan phase via a flag"). When a new SDLC-cycle variant is needed, add a separate composite toframework/composites.yaml(asship-tasksits next toship) — that is the only supported model.
To add a new composite: write framework/atoms/<name>.md files for any new phases (or extend an existing atom with a new <param-branch> value), write a framework/composites/<name>.md wrapper (frontmatter + body with {{PHASES}} marker + <inline-phase index="N"> and <gate after="N"> blocks as needed), add the composite entry to framework/composites.yaml, then run the generator and deno task check.
Benchmark Fixture deno.json Contract
When an acceptance test scenario invokes deno task check (or any project-check command) inside the sandbox, the fixture's deno.json MUST exclude the copied framework + bench artefacts from fmt/lint/test:
"fmt": { "exclude": [".claude/", "documents/", "acceptance-tests/"] },
"lint": { "exclude": [".claude/", "documents/", "acceptance-tests/"] },
"test": { "exclude": [".claude/", "acceptance-tests/"] }
Without these, the sandbox's deno fmt --check and deno lint apply to the copied framework/-as-.claude/ tree, which has formatting/lint drift relative to the sandbox project. Symptoms: a happy-path checklist item like check_gate_enforced fails with the agent surfacing an fmt diff against files it never touched. Fix-it-once template lives in framework/core/skills/implement/acceptance-tests/tdd-cycle-completes/fixture/deno.json.
Framework primitive placement
When a task creates a new framework primitive, decide the subdir FIRST:
- User-invoked via
/<name>(no model auto-discovery) →framework/<pack>/commands/with short kebab-case names. Examples:/commit,/update,/review-and-commit. - Model auto-invocable (skill activation by description match) →
framework/<pack>/skills/with short kebab-case names. Examples:deep-research,draw-mermaid-diagrams.
Picking the wrong subdir fails check-naming-prefix.ts (NP-3) and requires a file move + SRS/SDS location edits. The CLI writer injects disable-model-invocation: true automatically for commands/ — do NOT set it in source.
Acceptance Test Infrastructure Smoke Test
Before writing or modifying a benchmark scenario for a command or skill, run one existing scenario for the same primitive to verify infrastructure works:
deno task acceptance-tests -f <existing-scenario-id>
If it finishes with 0 agent steps or "Unknown skill" — the acceptance test runner has an infrastructure bug (e.g., copyFrameworkToIdeDir not copying the primitive). Fix the runner first; do not write new scenarios on broken infrastructure.
The runner also pre-checks that scenario.skill is mounted in the sandbox before spawning the agent and warns on suspiciously short agent output (< 200 chars with exit 0).
- The
acceptance-tests -fflag accepts ONE substring (last-wins on multiple). To run several scenarios: use a broader substring covering all of them, OR run sequential single--finvocations. Multiple-fflags silently keep only the last value. - An
acceptance-testsrun reporting "0 errors, 0 scenarios run" with exit 0 is a SETUP FAILURE, not success. Check stderr for "Error running scenario" lines. Common cause: missingfixture/directory referenced by the scenario's setup hook. - A scratch worktree starts with an empty
.acp-bridges, and every scenario in it dies with exit 124. The runner resolves the ACP bridge from the tree it runs in, so a worktree made for a parent baseline or a RED probe has none: the agent never starts, the judge scores an empty transcript, and the timeout reads as "the agent hung". The log line names it —ACP bridge @agentclientprotocol/codex-acp@1.1.7 is not installed in <worktree>/.acp-bridges. Rundeno task acp-bridgesinside the worktree before the first scenario, not after the first failure (observed 2026-09-19, one aborted run). - Read the duration and token counters before the verdict. A scenario that exits 0 after ~17 s and ~25k tokens did not measure behaviour — the session was cut short, and the judge then scores whatever fragment exists (2026-09-19: an "interrupted inspection attempt" scored as a result). Compare against a normal run of the same scenario: these run 60–200 s and 200k–1.2M tokens. On a short run, clear the scenario's entry under
acceptance-tests/cache/and re-run; the cache will otherwise hand you the truncated verdict again. 0 agent stepsis not an auth failure until you have checked auth in a REAL environment.claude auth statusrun underenv -ireports"loggedIn": falseon a fully authorised machine — the stripped environment cannot reach the credential store. Re-run it with the inherited environment before concluding anything, and read the Keychain entry'smdattimestamp (security find-generic-password -s "Claude Code-credentials") to see when the token was last refreshed. Asking the user to log in again on the strength of anenv -iprobe spends a round-trip on a state that was already fine.- user-level skill collision (FR-ACCEPT-ISOLATION): Claude Code's Skill tool resolves
~/.claude/skills/<name>/SKILL.md(user-level) over<sandbox>/.claude/skills/<name>/SKILL.md(project-level) on name collision. Without mitigation, every framework-sourceSKILL.mdedit silently routes the model to the developer's installed snapshot, and the Acceptance Test TDD RED→GREEN cycle produces no observable change. Mitigation lives inprepareAcpClaudeHome(scripts/acceptance-tests/lib/acp/auth.ts, wired into the Claude profile'sprepareWorkspace; the directClaudeAdapterwas retired with the ACP migration): builds<workDir>/bench-home/(sibling of the sandbox; placed outside the sandbox cwd sogit statusdoes not flag it as untracked) with an empty.claude/skills/and symlinks back to~/Library/Keychainsand~/.local/share/claudefor OAuth/Keychain auth, then exportsHOME=<workDir>/bench-home.~/.claude/skills/is never read or written by the bench. Cursor/Codex/OpenCode have no analogous bug and pass through unchanged.