Imported from qaml-ai/camelAI (
AGENTS.md). Install upstream withnpx skills add qaml-ai/camelAI. Copyright stays with the author.
camelAI Agent Guide
Keep this file concise and durable. Add details here only when they help future agents navigate the repo or avoid common mistakes. Feature-specific behavior should usually live near the code or in tests.
What This Is
camelAI is an AI coding assistant platform on Cloudflare Workers + Durable Objects with Cloudflare sandbox containers for builds/analysis. Users chat with persistent coding workspaces; each thread runs camelAI's pi-based coding harness in ChatThreadDO, and generated apps publish to *.camelai.app / environment-specific app hosts.
High-Level Architecture
React Router SSR + browser WS
|
v
Cloudflare main Worker + Durable Objects
|
v
Project files in WorkspaceFilesystemDO + R2 (do-r2 backend)
Builds/deploys + analysis in Cloudflare sandbox containers
Dispatcher Worker routes published user apps.
R2 stores uploads/assets/previews.
Cloudflare AI Gateway and BYOK credentials back model access.
Agent turns run in ChatThreadDO (Pi coding agent). Project source files live
in WorkspaceFilesystemDO + R2 (every project is backend: "do-r2"); builds,
deploys, and analysis run in Cloudflare sandbox containers (ProjectBuildContainer,
AnalysisContainer, DbQueryContainer). The legacy Azure project-runtime-service VM
and its PROJECT_RUNTIME_HOST bridge are gone; the only remaining VM is the
static-IP database egress relay (infra/db-egress-relay/, see docs/db-egress-relay.md).
SQL queries/exports run in DbQueryContainer (the DATA_PROXY binding is served
worker-side). There is no in-repo Go sandbox-host or data-proxy tree.
Repository Map
src/- React Router 7 app, routes, loaders/actions, UI components, shared server/client libraries.src/routes.ts- Imperative React Router route config. Add page/API routes here.src/routes/api/- React Router API routes for most user-facing REST (billing checkout, workspaces, chat groups, etc.).src/components/ui/- shadcn/ui components.workers/main/- Main Cloudflare Worker, Durable Objects, HTTP/SSE transports, admin MCP, admin APIs, proxies, container image Dockerfiles.workers/main/src/identity/-UserDO/OrgDOand related identity helpers (auth.tsis a compatibility barrel).workers/main/src/routes/- Worker-native HTTP (SSE streams, Stripe webhook, admin MCP, most/api/admin/*on Hono). Prefer documenting new paths here vssrc/routes/api/— see API routing below.workers/dispatcher/- Workers for Platforms dispatcher for deployed user apps.workers/app-usage-guard/- Account-wide Durable Object SQLite usage monitor and reversible app quarantine Worker; seedocs/deployed-app-usage-guard-design.md.workers/user-logs-tail/- Tail worker for deployed app logs.workers/e2e-reports/- Public viewer ate2e-reports.camelai.devserving Playwright E2E reports from R2 (uploaded by the E2E workflow); deploy withbun run deploy:e2e-reports.workers/eval-reports/- Read-only results store + viewer for agent evals atevals.camelai.dev(evals run locally;EVAL_REPORT=1publishes them); deploy withbun run deploy:eval-reports.- The Go data-proxy (external
qaml-ai/project-runtime-servicecmd/data-proxy) is retired: SQL queries and warehouse exports now run in theDbQueryContainerCloudflare container (workers/main/src/db-query-service.ts+data-proxy.tscompat surface), and theSANDBOX_HOSTVPC binding is gone. Do not reintroduce either. Decommission checklist:docs/db-egress-relay.md. sandbox/- Agent skills, project scaffold templates (create-worker/), and the canonicalvalidate-notebook.py(byte-copied intoworkers/main/analysis-sandbox-assets/for the analysis image build context). Not the agent control plane or harness — those live inworkers/main(chat-thread-do.ts, Pi tools, Dockerfiles).scripts/- Deploy, eval, self-host, and maintenance scripts.docs/- Supporting documentation; seedocs/README.mdfor the canonical index (many*-plan.md/ feedback files are historical).plans/- Active cross-cutting architecture plans (e.g. OrgDO split, no-VM build/deploy).infra/- Terraform for the static-IP database egress relay VM (infra/db-egress-relay/);infra/selfhost/for self-host cloud templates. Seeinfra/README.md.tests/- Vitest UI /src/libunit tests (vitest.config.ts).workers/main/tests/- Worker / Durable Object / Miniflare tests +evals/(vitest.workers.config.ts).e2e/- Playwright end-to-end specs..agents/skills/- Agent skills for this repo (evals and shadcn).
API routing
Two HTTP surfaces share the main worker:
| Surface | Location | Typical contents |
|---|---|---|
| React Router | src/routes/api/ |
Session-cookie user REST (workspaces, billing checkout, uploads, chat groups) |
| Worker-native | workers/main/src/routes/ |
SSE streams, Stripe webhook, admin MCP, most bearer admin REST |
workers/main/src/index.ts routes some paths (e.g. /api/admin/*) to worker modules before React Router SSR. When adding an API, match an existing neighbor; do not invent a third pattern.
Internal vs product names
Product name is camelAI. Internal Cloudflare resources, DO/MCP class names, headers, and Analytics Engine datasets often still use legacy internal codenames (for example the *_observability_* / *_errors_* dataset prefixes). Prefer the existing name in code; do not rename bindings or exported DO classes casually.
Development Commands
Use Bun for JS commands.
bun run dev # React Router dev with Cloudflare bindings, default localhost:3001
bun run build # Production React Router build
bun run typecheck # Generate route types, then tsc
bun run lint # ESLint
bun run test # Vitest watch mode
bun run test:run # Vitest run once
bun run test:workers # Worker/Miniflare tests
bun run test:all # Unit + worker tests
bun run test:e2e # Playwright
The bundled shadcn/ui catalog served by add_shadcn_component and the scaffold's seeded
primitives live in workers/main/src/shadcn-registry.generated.ts; refresh it from the public
registry with bun scripts/generate-shadcn-registry.mjs (also update
workers/main/project-build-sandbox-warmup/package.json — tests/project-scaffold-warmup.test.ts
enforces the sync so the build-sandbox image keeps all installable packages in its bun cache).
Common deploy commands:
bun run deploy:main:prod
bun run deploy:main:staging
bun run deploy:dispatcher:prod
bun run deploy:dispatcher:staging
bun run deploy:dispatcher:evals # testing-grounds dispatcher for real-deploy evals
bun run deploy:usage-guard:prod
bun run deploy:usage-guard:staging
Real-deploy evals (testing grounds)
Agent evals deploy apps for real to a dedicated testing-grounds namespace so they are actually
usable. The agent deploys with the normal deploy_project tool, which uploads via
deployWorkerModulesDirect (direct-dispatch-deploy.ts) straight to the Cloudflare API from
the worker, so nothing eval-specific sits in the deploy path. It publishes to the
chiridion-platform-evals dispatch namespace and registers in OrgDO exactly like production —
so list_apps / set_preview and AgentEvalSessionResult.deployedApps surface the app through
the normal app path with no eval-specific branches in chat-thread-do. The testing-grounds host
comes from the eval env's WORKER_BASE_URL / LOCAL_APP_VANITY_DOMAIN (*.evals.camelai.app),
and virtual bindings resolve against the staging main worker (CF_WORKER_NAME); these are
pinned in wrangler.test.jsonc. Real deploy is the default for agent eval runs whenever
CF_API_TOKEN is set; EVAL_REAL_DEPLOY=0 disables it (deploy evals then skip; the gate is
isRealEvalDeployEnabled in eval-deploy-context.ts). Served by the evals dispatcher
(workers/dispatcher/wrangler.evals.jsonc); the namespace + DNS routes are created out-of-band.
Eval apps are kept (no cleanup). Live-data bindings (DATA_PROXY/CONNECTIONS) won't resolve
to the eval's local workspace; self-contained apps render fully.
Eval results viewer (workers/eval-reports/)
Agent evals run locally (they need Docker + .dev.vars; Miniflare spawns the eval sandbox
containers via the local Docker daemon): bun run test:eval <id> (or the :dashboard / :deploy
/ :sandbox shortcuts) wraps scripts/run-agent-eval.mjs; scripts/run-eval-suite.sh runs a
list/all. Every eval's thread runs on a local agent runtime (Docker Compose from
qaml-ai/run's deploy/selfhost with its dev override; node scripts/runtime-eval-harness.mjs up|run|status|down, started by run-agent-eval.mjs when it is not up): the tests call
runRuntimeEval (workers/main/tests/evals/runtime-eval.ts), and the runtime reaches chiridion's
/mcp/agent through the eval relay (scripts/lib/eval-runtime-relay.mjs). Needs ~/agent-runtime
(AGENT_RUNTIME_DIR) or AGENT_RUNTIME_IMAGE; see the running-agent-evals skill. There is no remote runner — the retired qaml-ai/camelai-eval-runner VM control plane
was replaced by a shared results store + read-only viewer on Cloudflare (workers/eval-reports/,
evals.camelai.dev: Worker + R2, everything behind Cloudflare Access, no worker secrets).
Set EVAL_REPORT=1 on a run to publish it there when it finishes: run-agent-eval.mjs captures
the output log and invokes scripts/report-eval-run.mjs, which uploads the transcript artifact +
log and posts metadata (auth via an Access service token in CF_ACCESS_CLIENT_ID/SECRET, or a
local cloudflared access login). Use the running-agent-evals skill for the run/report/read
workflows; the always-current API doc is served at GET evals.camelai.dev/skill. Deploy the
viewer with bun run deploy:eval-reports; see workers/eval-reports/README.md.
Suite and matrix runs share an EVAL_BATCH_ID, and the viewer groups those reported runs into
batches.
Adding a new eval. workers/main/tests/evals/manifest.json is the single source of truth for the
committed eval list. To add one: (1) create workers/main/tests/evals/<id>.test.ts, gated on
RUN_AGENT_EVALS === "1", ending in emitEvalTranscript({...}) from ./eval-transcript (every eval
shares one transcript marker pair); (2) add a { "id", "description", "kind", "realDeploy"? } entry to
manifest.json. It is then runnable via bun run test:eval <id> and included in EVAL_TARGET=all —
no other files to edit. custom-prompt-live is intentionally not in the manifest (it's the generic
CUSTOM_EVAL_*-driven harness). run-agent-eval.mjs runs exactly one eval; run-eval-suite.sh
iterates a comma-separated list or all, running each eval in turn.
Analysis and builds do not run under the worker test pool yet. AnalysisContainer and
ProjectBuildContainer start named images from ctx.container.images (the durable_object
container policy), which the miniflare/workerd bundled with @cloudflare/vitest-pool-workers
predates, so vitest.workers.config.ts attaches no container to them and evals that run notebooks
or builds need a pool upgrade first. For local runs of the analysis image (self-host smokes, dev),
scripts/build-analysis-sandbox-image.mjs builds it natively for the host arch (on Apple Silicon
the Jupyter kernel never answers its handshake under Rosetta/QEMU; only the static sandbox-shim
is amd64) and rebuilds when the Dockerfile or the baked assets under
workers/main/analysis-sandbox-assets/ change (a content hash is stamped on the image as a label).
Container egress workaround (workerd#6793). Evals run the agent in a Cloudflare Container via
@cloudflare/vitest-pool-workers/Miniflare. On newer hosts (Linux kernel ~6.17 / Docker 29.x) the
stock cloudflare/proxy-everything egress sidecar's TPROXY rules intercept docker bridge control
traffic, so the container never becomes ready and evals fail with kj/timer ... operation timed out
/ "Container failed to start". workers/main/eval-egress-fix/ is a thin wrapper image that adds a
bridge-bypass rule; run-eval-suite.sh builds it and run-agent-eval.mjs auto-selects it via
MINIFLARE_CONTAINER_EGRESS_IMAGE (both no-ops where the bug doesn't trigger). Such hosts also need
the docker bridge allowed to reach the host (e.g. ufw allow in on docker0). Remove the wrapper
once cloudflare/workerd#6794 ships in a release.
Separately, vitest-pool-workers leaves the eval container + sidecar running after each run
(workers-sdk#14242); they accumulate and exhaust the host. run-agent-eval.mjs prunes leftover
AnalysisContainer/ProjectBuildContainer containers before/after each run. The sweep is global,
so it's only safe when one eval runs at a time (the normal local case); an orchestrator that
runs evals concurrently must set EVAL_MANAGED_CLEANUP=1 to skip it and own cleanup itself.
Frontend Conventions
- React Router is in framework mode. Prefer
loader,action,<Form>, anduseFetcherover client-only fetching inuseEffect. - Route definitions live in
src/routes.ts; route modules live insrc/routes/. - Tailwind CSS v4 and shadcn/ui are the default UI stack.
- For UI work, use the
shadcn-componentsskill and existing primitives insrc/components/ui/. - Use
cn()from@/lib/utilsfor class composition. - Use Lucide icons where appropriate.
- Keep app surfaces work-focused and dense. Avoid marketing-style pages unless the task explicitly asks for one.
Worker And Durable Object Conventions
Important DOs and runtime classes live primarily in workers/main/src/:
identity/(auth.tsbarrel) -UserDO,OrgDO, and identity helpers. OrgDO domain extraction is in progress (plans/split-auth-durable-objects.md); prefer new org logic inidentity/org/modules rather than growingorg-do.ts.workspace.ts-WorkspaceDO, workspace metadata, integration state, token refresh alarms.chat-thread-do.ts-ChatThreadDOcompatibility façade and chat transport/turn orchestration (the shared connection bridge ischat-thread/sse-connection.ts, with WebSocket and HTTP polling sinks). Focused collaborators live inchat-thread/(Pi persistence, model/tool setup, UI mirroring, recovery journals, verified completion evidence, metadata, preview/access/automation, and streaming activity); verify this surface withbun run test:workers -- chat-thread.workspace-cron.ts-WorkspaceCronDO, scheduled prompt storage and dispatch.worker-logs-do.ts-WorkerLogsDO, recent deployed-app logs written by the tail worker and read over RPC (in-memory ring buffer; not SQLite-persisted).admin-index-do.ts-AdminIndexDO, admin indexes and dashboard-style aggregates.email-handle-registry.ts-EmailHandleDO, email handle ownership.*-mcp.ts/connections-runtime.ts- Per-provider connection MCP wrappers and shared connection runtime (candidate for anintegrations/folder).observability.ts- Shared Cloudflare Analytics Engine event/error writer. New structured instrumentation should go through this helper instead of callingwriteDataPointdirectly.lake-streams.ts- Tool-call telemetry export (one row per CodeModeToolsBinding call) to an Iceberg table in R2 Data Catalog via Cloudflare Pipelines. The binding is optional and the helper no-ops without it. Setup and queries:config/pipelines/README.md.
Durable Objects use SQLite-backed storage. Prefer:
this.ctx.storage.sql.exec("SELECT * FROM table WHERE id = ?", id);
this.ctx.storage.kv.put("key", value);
const value = this.ctx.storage.kv.get("key");
Do not use legacy async DO storage (await ctx.storage.get/put) in new code. Do not use module-level mutable Map, Set, or singleton instance caches in Worker code; isolate reuse can leak stale state across requests.
For background work in Workers, import waitUntil from cloudflare:workers and catch/log failures:
import { waitUntil } from "cloudflare:workers";
waitUntil(
task().catch((error) => console.error("Background task failed", error)),
);
Observability
- Cloudflare Workers Observability and source-map uploads are enabled in deployed Wrangler configs.
- Structured operational events go to
OBSERVABILITY_EVENTS; structured errors are mirrored throughERROR_ANALYTICS. UserecordObservabilityEvent/recordErrorEventfromworkers/main/src/observability.tsfor new instrumentation. - Keep observability payloads diagnostic but not transcript-like: include ids, counts, status, durations, routes, and error metadata; do not store chat message contents, secrets, request bodies, or auth headers. The one deliberate exception is the transcript data lake (
config/pipelines/README.md), which exports transcript text on purpose and is governed by its own access/retention rules — it is not a licence to widen Analytics Engine payloads. - The main app workers attach Tail Consumers to
workers/user-logs-tail/, which forwards raw Worker trace/log/exception events intoWorkerLogsDO. - Production datasets are
chiridion_observability_prodandchiridion_errors_prod; staging uses the corresponding_stagingdatasets. Verify bindings in the environment-specificwrangler*.jsoncfiles before changing collection paths. - Query Analytics Engine through Cloudflare's SQL API with an account token that has Account Analytics Read. The account id is
CF_ACCOUNT_IDin Wrangler vars. Example:
curl "https://api.cloudflare.com/client/v4/accounts/$CF_ACCOUNT_ID/analytics_engine/sql" \
--header "Authorization: Bearer $CF_API_TOKEN" \
--data "SELECT timestamp, blob1 AS event, blob3 AS component, blob5 AS status, blob9 AS thread_id, double2 AS duration_ms FROM chiridion_observability_staging WHERE timestamp > NOW() - INTERVAL '1' HOUR ORDER BY timestamp DESC LIMIT 100 FORMAT JSON"
OBSERVABILITY_EVENTSschema:blob1 event,blob2 severity,blob3 component,blob4 operation,blob5 status,blob6 route,blob7 method,blob8 path,blob9 threadId,blob10 workspaceId,blob11 orgId,blob12 userId,blob13 requestId,blob14 provider,blob15 model,blob16 errorName,blob17 errorMessage,blob18 errorStack;double1 timestamp_ms,double2 duration_ms,double3 status_code,double4 count,double5 size;index1 sample key. A few events carry extra numeric dimensions (extraCounts) appended fromdouble6on, so the fixed positions above never move; each such event documents its own order —pi_context_budgetisdouble6 image_count,double7 image_chars,double8 message_count,double9 result_bytes(payload bytes of the view that actually shipped), withdouble4the estimated context tokens anddouble5the estimated payload bytes going in. Itsblob5 statusis one ofunchanged/memo_hit/row_hit/summarized/no_cut; more than onesummarizedper turn is a regression, and anyno_cutmeans an over-budget context shipped whole.ERROR_ANALYTICShas the error-focused subset:blob1 event,blob2 component,blob3 operation,blob4 status,blob5 errorName,blob6 errorMessage,blob7 threadId,blob8 workspaceId,blob9 orgId,blob10 userId,blob11 requestId,blob12 route,blob13 path,blob14 errorStack; doubles match the same timestamp/duration/status/count/size order.- Runtime-thread events (
OBSERVABILITY_EVENTSpositions;blob9-11thread/workspace/org where known):runtime_thread_send_failed: componentruntime_thread,blob4send/first_send,blob5refused(blob16the refusal code, e.g.usage_limit) /busy/runtime_4xx/runtime_5xx/exception,double3the runtime's HTTP status. Thrown ones also go toERROR_ANALYTICS.runtime_token_mint_failed: componentruntime_thread,blob4token_route/page_seed(the page's first load, token + history),blob5no_agent(404: the thread has no agent yet) or a thrown error's class as above,double3status.runtime_watch_error(browser, via /api/client-errors): componentbrowser,blob4runtime_watch,blob5the HTTP status or error name; its details carryphasestart/watch.runtime_event_handler_failed(alsoERROR_ANALYTICS) /runtime_event_unknown_thread: componentagent_runtime_events,blob4the event type,blob13the event id. A handler failure answers 500 and the runtime redelivers.thread_running_lease_expired: componentworkspace_do,blob4lease_sweep,blob5any_backend(the running row does not record which backend ran the turn),double2ms since its last heartbeat.thread_running_lease_refresh_missed:blob4lease_refresh,blob5the refresher (runtime_usage,chat_thread_do); a refresh that found no running row (a late tick, or a lease that expired mid-turn).agent_mcp_call: componentagent_mcp, one per tool call on/mcp/agent;blob4the tool name,blob5ok/error/forbidden(authorization refused) /input_required/exception(thrown;blob16its error name),blob9–blob12thread/workspace/org/user,double2duration ms. No arguments or output.agent_mcp_auth_failed: the runtime token was refused (401),blob5no_token/invalid_token(expired, wrong audience or issuer, another tenant).
- For aggregate counts/sums, account for sampling with
_sample_interval, for exampleSUM(_sample_interval)instead ofCOUNT().
Chat And Runtime Flow
-
Chat uses native WebSockets with immediate HTTP polling fallback on socket errors or unexpected closes (including open-then-close). There is no added WebSocket connection timeout. Protocol, limits and tests:
workers/main/src/chat-thread/transport.md. -
Browser chat transport is native WebSocket at
/agents/chat-thread/:threadId, including RPCs and resume frames. A socket error or unexpected close immediately switches that view to completed HTTP polling responses atGET /agents/chat-thread/:threadId/sse?transport=poll; sends then usePOST /agents/chat-thread/:threadId/call. The old SSE receive mode remains server-side for already-open older bundles.SseAgentClient/useSseAgentretain their historical export names but new browser connections use WebSockets. Explicit policy denials stay terminal; unmounts do not trigger fallback. Workspace status remains SSE at/api/workspaces/:id/status/stream; the old workspace socket remains removed. Client opens arechat_ws_openorchat_poll_open; existing error telemetry names remain for continuity. -
The main worker validates access (
authorizeChatTransportRequest) and routes toChatThreadDO, which bridges WebSocket and HTTP receive sessions into the partyserver connection model via a syntheticSseConnection(workers/main/src/chat-thread/sse-connection.ts) — the wrappedonConnect/onMessage/onClosechains and the resume handshake run unchanged. Design + invariants:plans/sse-migration/DESIGN.md(untracked, kept in the repo checkout). -
ChatThreadDOruns the Pi coding agent in the Durable Object. Project file operations go toWorkspaceFilesystemDO+ R2 (do-r2backend); builds/deploys/analysis run in Cloudflare sandbox containers. There is no shell/bashtool — the agent uses the DO-backed file tools plusdeploy_project/add_dependencyandjs_exec.deploy_projectbuilds, publishes, returns the live URL, and opens preview; no manualset_previewis needed, though the tool remains available for explicit preview switches.run_notebooklikewise opens a clean successful notebook run in preview automatically and leaves preview unchanged on failure.dry_run: truevalidates a deploy without publishing. -
Transport + render history are owned by
@cloudflare/ai-chat(ChatThreadDO extends AIChatAgent). A turn is a resumable UIMessage stream:onChatMessageruns the Pi prompt (or the recovery/resume branch) and relays Pi runtime events through the encoder as native UIMessage chunks;chatRecoveryowns bounded re-drives of an interrupted turn andchatStreamStallTimeoutMsbounds a stalled turn (its stream-cancel disposes the hung Pi session and routes the turn into recovery).chatRecovery's budget is PROGRESS-GATED, so a turn that journals a checkpoint and then kills the isolate every pass renews it forever: thepiActiveTurnmarker additionally carries progress-independent re-drive counters that abandon such a turn — commit the journal tail, teardown, durable terminal — instead of resuming it. They are split by cause:isolateDeathResumeAttempts(PI_TURN_RESUME_BUDGET, charged only when nothing in-process observed the interruption),voluntaryResumeAttempts(PI_TURN_VOLUNTARY_RESUME_BUDGET, for transient-provider-retry / config-change re-drives) and a loose total (PI_TURN_TOTAL_RESUME_BUDGET), so ordinary deploy resets and provider 529s cannot abandon a healthy turn. The isolate-death count also picks a RECOVERY LADDER rung (piTurnResumeRung), each cheaper in memory than the last: deaths 1-2 resume normally, the 3rd resumes DEGRADED (eager compaction + a hard image-hydration budget, applied per provider request viatransformPiProviderContext; the compaction is EPHEMERAL — nopi_core_compactionrow — and floored atPI_DEGRADED_COMPACTION_FLOOR_FRACTIONof the real threshold, so a recovery can never permanently truncate a thread), the 4th skips the provider entirely and SALVAGES the journal (settled work + an "ask me to continue" note committed as the final assistant message, turn closed out normally), and anything past that is the terminal abandonment. The client renders fromuseAgentChat(resume: true); no bespoke websocket transcript fan-out. -
Two message stores:
pi_core_*tables are Pi's model-side transcript (authoritative for the agent and repair/eval tooling); the ai-chat message table is the browser render history. A high-water-mark backfill (topUpUiMessagesFromPiCore/getUiMessagesRPC) mirrors new pi_core rows into render history. Same-content-same-id invariant: every pi_core row is stamped with the render message id it streams into (uiMetadata.renderMessageId— turnId for assistant rows, the persisted skeleton's id for user rows), so the mirror is an idempotent upsert and a whole turn folds into the one live message id. The chat-route loader callsgetUiMessagesfor cold load only; live sync rides the stream + CHAT_MESSAGES broadcasts. Seedocs/chat-transcript-simplification.mdfor the invariants and the derive-on-read roadmap. -
Settled render history is derived AT THE STORAGE BOUNDARY (
workers/main/src/chat-thread/derived-render-page.ts): pi_core rows are paged newest-first byidx(metadata first, then payloads one at a time), derived incrementally, and the walk stops once the 50-message / 4MB window is covered. On that SETTLED READ path peak memory is bounded by the page (a window admits at mostCHAT_RENDER_WINDOW_MAX_BYTES * PI_DERIVE_MAX_WINDOW_BYTE_FACTOR= 8MB of stored payload plus one oversized row), not by the thread. It is NOT an end-to-end O(page) claim: the compaction WRITE path (materializeSettledRenderArchiveFromPiCore) and the mirror rebuild (ChatThreadUiMirror.topUpUiMessagesFromPiCore, reachable only viauiRender: "rebuild"/ admin resync / fork seeding) still load the whole transcript, as does the pi session's own provider context — see open_issues. It replaced a path that materialized the whole transcript plus the whole ai-chat archive table before paginating, which OOM-killed the DO on every load of a 5,232-row thread. Two invariants govern it: a page NEVER cuts arenderMessageIdfold below the window byte ceiling (a turn is atomic for pagination); past that ceiling the window closes mid-fold, reportsfoldCutsat warn severity, and the two halves are reunited byprependOlderRenderMessages, which MERGES parts for a duplicate id instead of dropping the arrival, and windowed rows reconstruct their absolute position in the legacy full load, because an unstamped user row's derived id embeds it. Older-page cursors aredp:p:<piRowIdx>while the derived tail lasts anddp:a:<aiChatChronologyCursor>once a page reaches back past it — the pre-compaction archive is paged lazily fromcreated_at— the seam is the MINIMUM pi timestamp in the derived window (nevermessages[0], which after a preserve-compaction is the wall-clock-stamped "[Context Summary]" row) and archive rows are additionally filtered against the derive's ids androle+createdAtMskeys, becausepi_user_<ts>_<index>ids renumber across a compaction — so an ordinary load never reads it (legacyd:<index>cursors are accepted and degrade to "serve the newest page").getDerivedUiMessagePageis the one seam; the golden pagination-equivalence test against the old full-thread pager isworkers/main/tests/chat-thread-derive-pagination.test.ts. -
The Pi event → UIMessage chunk encoder is
src/lib/pi-chunk-encoder.ts; the UIMessage → legacyMessagerender adapter (both directions) issrc/lib/ui-message-adapter.ts. Steering appends via RPC +persistMessages. The Agent-state payload (src/lib/chat-agent-state.ts, shared DO/client) now carries only coarse fields (preview, todos, title, model, terminal error) — streaming and turn duration/completion are derived from the hook +message-metadata.pi. -
Client seed/stream ownership seam: on a tab switch, the snapshot-derived
useAgentChatseed EXCLUDES the mid-stream assistant message (resolveDisplayChatData); the resumed stream rebuilds it from scratch and Chat bridges the paint gap (bridgedStreamingMessageId). Don't reintroduce hydrated in-flight content into the seed — replay onto it duplicates parts (upstreamai/agentsmerge is replace-last-or-push). -
Thread records store provider/model state on org thread data. Verify current fields in
OrgDObefore changing related behavior. -
Slash commands are allowlisted in
ChatThreadDO; checkSLASH_COMMANDSbefore adding or changing one. -
Clarifying questions use the Pi
AskUserQuestion/ask_user_questiontools. -
Hosted agent runtime migration (in progress,
plans/agent-runtime-migration.md): when a deployment has a runtime tenant (AGENT_RUNTIME_API_TOKEN,AGENT_RUNTIME_TENANT,AGENT_RUNTIME_DEFINITION), each new thread whose model has a runtime route (agent-runtime/model-routes.ts) is pinned to the hosted runtime (agent_backend_pinnedrecords the backend and any fallback reason); other threads keep the in-DO loop.chat-thread/runtime-agent.tsstands in for the Pi Agent and relays the runtime's Pi events;/mcp/agent(routes/agent-mcp.ts) servesCodeModeToolsBindingtools; the runtime calls providers with key scopes (agent-runtime/key-scopes.ts) and reports usage asusage.recordedevents to/agent-runtime/events(routes/agent-runtime-events.ts,agent-runtime/usage.ts); only Codex goes through chiridion (/agent-runtime/llm/openai-codex/*,agent-runtime/codex-forwarder.ts). Tests:bun run test:workers -- agent-mcp agent-runtime code-mode-capability-tools.
Adding a new chat model
When adding a model from Anthropic, OpenAI, OpenRouter, or another provider,
follow the checklist at the top of src/lib/model-catalog.ts. The picker,
pricing, and harness routing live in separate files, and the catalog tests fail
if any of them drift apart.
Uploads, Files, And Safety
- Chat uploads use multipart R2 upload APIs under
/api/workspaces/:id/upload. - Workspace file API routes live under
/api/workspaces/:id/fs/*. - File safety logic lives in
workers/main/src/file-safety.tsand is applied before agent turns for suspicious uploaded-file/deploy/bridge workflows. - The Pi system prompt is assembled in
workers/main/src/chat-thread-do.ts; keep security-relevant prompt changes explicit and tested.
Proxies And Bindings
- Sandbox containers do not get a generic Worker API proxy or any header-authenticated Worker routes. Container access to Worker services goes through DO-side outbound handlers that attach scope (for example the analysis sandbox's
connections.internal), and deploys go through the platform's own deploy tools (deployWorkerModulesDirectindirect-dispatch-deploy.ts), not through the container. - BYOK credentials are scoped by org/thread and should not be placed into container environment variables.
- User app deploys can rewrite internal service bindings such as the data proxy, virtual AI binding, and virtual R2 bucket. Relevant files include
workers/main/src/cf-api-proxy.ts,data-proxy-service.ts,ai-virtual-binding.ts, andr2-virtual-bucket.ts. - Outbound database traffic egresses from the sandbox host VM IP
20.46.233.68(surfaced in direct database connection setup UIs for firewall/VPC allowlisting; constant insrc/lib/sandbox-network.ts). DbQueryContainer(bound asDB_QUERY_SANDBOX; native Durable Object container, Sandbox SDK 1.0, no user code) is THE SQL query/export path — the connection MCP and theDATA_PROXYuser-app binding both go through the legacy-contract surface inworkers/main/src/data-proxy.ts→db-query-compat.ts→db-query-service.ts. It keeps the static-IP guarantee by dialing databases through a SOCKS relay on the sandbox host VM (infra/db-egress-relay/; design + smoke + decommission checklist indocs/db-egress-relay.md); with no relay configured it dials from the container's own IP. The query logic is shipped from the worker per call (not baked): the runnerworkers/main/db-query-sandbox-assets/runner/db-query-runner.mjsis embedded through thevirtual:db-query-runner-sourcealias (Vite?rawfor the main worker, WranglerTextfor dispatchers) and piped into node over stdin in one stateless run (DbQueryContainer.runRunner); exports write Parquet straight into the workspace's warehouse R2 prefix (anS3Mountover an R2 API token on Cloudflare; a copy into theWAREHOUSE_EXPORT_BUCKETbinding after the run on self-host). The relay forwarder (cloudflared access tcp) is a background process started under flock (startRelayForwarder); the container hasenableInternetand no outbound interception. Keep the SSRF denylists in that runner andinfra/db-egress-relay/gost.yaml.examplein sync.
Stripe Billing And Credits
- Org billing state lives on
org_infoJSON. Key fields includebilling_status, Stripe customer/subscription ids, purchased credit cents, included/granted credit cents, trial credit grant metadata, and the last included-credit invoice id. - Hosted model access is enforced in the Worker/DO inference path. Hosted
trialingandactiveusage requires positive included/purchased credits; BYOK can be used from the free onboarding path and does not consume camelAI credits;enterprisebypasses Stripe subscription and credits. - Hosted credit allowances come from
src/lib/billing-plans.ts: Starter includes $10/month, Pro includes $40/month, and Team includes $50/month per paid seat.BILLING_TRIAL_CREDIT_CENTSandBILLING_SUBSCRIPTION_INCLUDED_CREDIT_CENTSare global emergency overrides; do not set them for normal tier-specific pricing. - Admins can grant credits manually with
POST /api/admin/orgs/:id/credits; credits add tobilling_credit_grant_total_centsand can use an idempotency key. STRIPE_MODEcan be set totestorlive; Stripe API calls reject secret keys whosesk_/rk_prefix does not match. Staging should useSTRIPE_MODE=test, and production should useSTRIPE_MODE=live.- Stripe webhooks land on
POST /api/billing/stripe/webhook. Subscription events sync status and grant the one-time trial cap;invoice.payment_succeededgrants recurring included credits idempotently; credit checkout sessions increment purchased credits. - Credit balance is purchased credits plus included/granted credits minus sandbox-host usage rows marked
credit_chargeable = 1.
Auth, Onboarding, And Admin
- Session/auth helpers are split between app-side loaders/actions in
src/lib/and Worker-side helpers inworkers/main/src/helpers/. - Reverse-proxy identity providers (auto-login behind Cloudflare Access or Pomerium) share one engine in
workers/main/src/helpers/proxy-auth-core.ts(JWT verify, JWKS, org mapping, revalidation); per-provider adapters areaccess-session.ts(RS256, get-identity endpoint) andpomerium-session.ts(ES256, inline-group claims). The registry/dispatcher isproxy-auth-providers.ts; app-side silent login/provisioning issrc/lib/proxy-auth.server.ts. To add a provider, implementProxyAuthProviderand register it. Tests:tests/{pomerium,cloudflare-access}-*.test.ts. Docs:docs/pomerium-auth.md,docs/cloudflare-access-auth.md. - Self-host Compose uses containerized Caddy as its ingress/TLS service. Bundled Pomerium is plaintext on loopback
127.0.0.1:5444; do not restore direct Pomerium TLS or a host-installed Caddy service. TLS modes areautomatic(Cloudflare/Route 53 DNS validation),external, andprovided. - Self-host agent customization (additive skills + prompt append/prepend) loads from
.selfhost/agent/at workerd-config generation; seeSELF_HOSTING.mdandscripts/selfhost-agent-pack.mjs. Verify withbun run test:workers -- selfhost-agent-packandbun run test:run -- selfhost-agent-pack-loader. - Password auth, OAuth account creation, email verification, onboarding, bans, and blocked signup policies all have tests in
workers/main/tests/; update or add focused tests when touching these flows. - First-touch marketing attribution and the durable first-accepted-message definition of
new_camel_activationare documented inMARKETING_ATTRIBUTION.md. - There is no superuser web UI; admin work goes through the admin REST API and admin MCP below.
- Bearer-auth admin APIs live under
/api/admin/*; implementation is inworkers/main/src/routes/admin/(entity detail/management endpoints inmanagement-routes.ts); the spec is served at/api/admin/openapi.json. - Admin MCP is served at
/api/admin/mcp(https://staging.camelai.dev/api/admin/mcpin staging) and uses OAuth scopeadmin:mcp. Staging is also behind Cloudflare Access; passCF-Access-Token: $(cloudflared access token -app=https://staging.camelai.dev)when connecting withmcporter. If an MCP client opens an authorize URL withscope=openid+email+profile, the flow will fail withinvalid_scope; forceadmin:mcpwithoauthScopeor a pre-registered static OAuth client. admin_js_execis the generic superuser remote Worker console for staging/production (binding RPC/fetch, Durable Objects, admin/self/outbound HTTP, assertions, and checked-in smoke suites). Seedocs/admin-js-exec.md; primitive env values and secrets are intentionally non-readable.- A reliable staging smoke path for admin MCP is: register or provide an OAuth client for the chosen localhost callback with
scope: "admin:mcp", setACCESS_TOKEN=$(cloudflared access token -app=https://staging.camelai.dev), then add a privatemcporterconfig entry withbaseUrl: "https://staging.camelai.dev/api/admin/mcp",auth: "oauth",oauthScope: "admin:mcp", andheaders: { "CF-Access-Token": "$env:ACCESS_TOKEN" }. Runnpx mcporter auth <server-name>followed bynpx mcporter list <server-name> --json. The browser session must be a camelAI superuser, otherwise authorization fails withAdmin access required. - Admin and moderation flows often involve durable tombstones in KV plus destructive cleanup. Avoid changing ordering without tests.
Integrations And Ingress
- Slack ingress starts in
workers/main/src/slack-events-queue.tsand routes turns intoChatThreadDO. - Email ingress starts in
workers/main/src/email-ingress.ts; workspace addresses are subaddressed by org/workspace slug. - Local Email Worker ingress can be simulated with
POST /cdn-cgi/handler/emailon the local dev server, passingfromandtoquery params plus a raw RFC 822-style body. Real MX-routed inbound email always reaches the deployed Worker route, not localhost. - Local outbound email uses the
send_emailbinding from Wrangler config. For agent email, sender addresses must resolve to workspace email handles onWORKSPACE_EMAIL_DOMAIN; do not fall back toEMAIL_FROM_ADDRESS/no-replyfor agent sends. - OAuth integration code is split across
workers/main/src/services/oauth.ts, route files, and workspace integration storage. Admin MCP OAuth is implemented separately inworkers/main/src/admin-mcp-oauth.ts. - Imported API definitions, typed operation policies, generic HTTP fallback, and GA4 behavior are documented in
docs/integrations-runtime.md; usedocs/connections-improvement-guide.mdfor the living UX, quality, safety, and evaluation strategy. - Scheduled prompts are owned by
WorkspaceCronDOand exposed through MCP tools.
Project Runtime
- Projects are DO+R2 backed (
backend: "do-r2"): metadata and source files live inWorkspaceFilesystemDO(ProjectFilesystemClientfor per-project files), with Cloudflare Artifacts git history. - Builds/deploys run in
ProjectBuildContainer(project-build-container.ts, nativectx.container, Sandbox SDK 1.0; commands inproject-build-commands.ts); notebook analysis inAnalysisContainer(analysis-container.ts, nativectx.container); SQL inDbQueryContainer. - The build container stops after its idle window and takes far longer to start again than the deploy retry ladder spans.
deploy_project/add_dependencytherefore runensureBuildSandboxReady(project-build-readiness.ts) before their first sandbox call — oneexec("true")probe when warm, a budgeted re-probe loop when cold — and the existing 5-attempt ladder stays as the guard for post-readiness blips (it re-arms the gate between attempts, under one shared budget). The gate and the ladder live inproject-build-readiness.ts; the adminproject-build-verifyroute drives the same pair throughrunWithProjectBuildReadiness, so an operator repro cannot fail on a container the user-facing path would have waited for. ProjectBuildContainer(nativectx.container) has no session layer, so the build path has no zombie handling: every command is its own process under GNUtimeout, a missing file reads asnull, a command stopped at its timeout comes backtimedOut: true, and a call that finds no running container throwsProjectBuildContainerUnavailableError(recognized byProjectBuildContainerUnavailableError.is(): across the DO RPC hop only the message survives, prefixed with the name; the next call starts a fresh container).AnalysisContainer(nativectx.container) has no session layer or zombie heal either: every command is its ownbash -cunder GNUtimeoutwith the stack's env passed explicitly (ANALYSIS_BASE_ENV:exec()sees no imageENVand starts in/), a command stopped at its timeout comes backtimedOut: true(the service turns that intoerror: "Command timed out after <ms>ms"plus asandbox_exec_timeoutevent; a program exiting 124 by itself stays an ordinary failure), and a container that stops under a command throwsANALYSIS_ENVIRONMENT_RESTARTED_MESSAGE(plusanalysis_container_stopped), never re-run.prepare(access)starts the container with the internet off and registers, in this order, the R2 mounts (sandbox-mounts.ts: S3Mount + S3Gateway with an R2 API token on Cloudflare, a copy through the binding on self-host),connections.internal→AnalysisConnectionsGateway(agent only; org/workspace in its props) and a catch-all →AnalysisEgress(PyPI over HTTPS for the agent, 520 for everything else). A hostname intercept registered after the catch-all gets nothing, so a live mount that needs repair on a running container replaces the container instead. The agent container is named<workspaceId>, the deployed-app oneapp-<workspaceId>(export mount only, no egress); the class refuses access that does not match its name.- A probe on a stopped build container blocks while it boots, so every probe carries its own deadline (
PROJECT_BUILD_PROBE_TIMEOUT_MS, belowPROJECT_BUILD_COLD_START_BUDGET_MS) and the probe command a container-side bound. - Permanent build-container startup failures never wait (
isProjectBuildPermanentStartupError: the build image missing from the deployment, or no container application for the class). Transient causes are the probe's own failures, the exec deadline,ProjectBuildContainerUnavailableError, and Durable Object errors the runtime marksretryable. - While a cold boot is in progress the tools stream
Build environment is starting…to the client viaChatThreadDO.streamToolProgress(anitem/commandExecution/outputDeltaon the parentjs_execcall), and stampbuildEnvironmentonto the result so the agent does not re-deploy into the same boot window. - When a build FINISHES the tools call
noteBuildSessionActivity, which stretches the container's inactivity timeout (setInactivityTimeout) toPROJECT_BUILD_ACTIVE_SESSION_WINDOW_MSso a second deploy in the same chat does not pay another cold boot. Thedurable_objectscheduling policy has nomax_instances: warm build containers count only toward the account's concurrent container limits, so widen the window with that in mind. - Every sandbox exec-class call is bounded CLIENT-side by
createSandboxExecDeadline(sandbox-exec-deadline.ts): op-class default/max (the same constants the services forward container-side) + an op-class IO overhead + a 15s marshalling grace. That covers the whole exec-class surface, including preludes — db-query's relay-forwarder probes and mount ensures run through the same deadline, since theirtimeoutis enforced container-side only. The container's owntimeoutstays primary and its error still wins inside the grace; the deadline only stops a wedged container from holding a turn to the 20-minutePI_TURN_TOOL_HARD_TIMEOUT. The analysis project legs size their IO overhead to the tree (analysisProjectIoOverheadMs) because materialize/persist is one RPC round trip per file. - The SDK exposes no way to cancel an in-flight
exec(thesignaloption is only checked before dispatch;killProcesscoversstartProcessonly), so nothing is killed container-side — re-check on SDK upgrade. Two consequences are load-bearing:deploy_project/add_dependencyshare ONE budget across the retry ladder and an EXHAUSTED budget is terminal (runrefuses to dispatch rather than starting a build it would abandon into the same per-project workdir; the ladder stops on it), and cold-boot waits and backoff sleeps are charged OUTSIDE that budget viadeadline.excludingso a container wake cannot eat the build's own time. - Stop interrupts a RUNNING tool:
requestStop→piSession.abort()→ the tool's abort signal, whichpi-toolspasses INTOkeepPiTurnToolProgressAliveWhile. The wrapper rejects withOperation abortedand releases the heartbeat immediately rather than waiting for the abandoned RPC (which keeps running until its own deadline). - Code-mode tool failures are recorded at the
callToolEnvelopeseam ascode_mode_project_tool_call_failed, for BOTH surfaces: a throw, and an operational failure returned as a value ({ success: false }/{ ok: false }). Theprovidercolumn carriesthrow/value. Value-shaped failures used to be invisible — a gated deploy failing every attempt showed nothing in telemetry. Two value shapes are deliberately NOT failures (toolValueFailureMessage): a SHELL OUTCOME (exitCode/stdout/stderrpresent —analysis_exec/run_code/run_notebook/add_python_dependencysetok:falsefor any non-zero exit of user code, whoseerrorfield is the container's raw stderr and must never reach ERROR_ANALYTICS), and acancelled: trueuser-declined confirmation. Events are deduped on tool+message and capped per binding instance (CODE_MODE_TOOL_FAILURE_EVENT_BUDGET); the value-path message is bounded (CODE_MODE_VALUE_FAILURE_MESSAGE_MAX) so it cannot inflate that key, and carries no fabricated stack. Arguments and program output are never logged. - Subagent tools (
Agent/Explore/Research/Oracle) getPI_SUBAGENT_ABORT_GRACE_MSinkeepPiTurnToolProgressAliveWhile: they CAN cancel, andchild.prompt()returns the accumulated answer afterchild.abort(), so a stop keeps that work instead of persisting an emptyOperation abortedresult. Sandbox-backed tools keep zero grace. - The legacy Azure
project-runtime-serviceVM, itsPROJECT_RUNTIME_HOSTbridge, the VMbash/vm_exec/clone_projecttools, and all their deploy/dev/migration scripts have been removed. The only remaining VM is the static-IP database egress relay (infra/db-egress-relay/). Do not reintroduce a project VM runtime.
Testing Guidance
- Place tests next to the surface they cover:
tests/— React Router UI,src/lib, and other non-worker unit tests (bun run test/test:run).workers/main/tests/— Worker, Durable Object, and Miniflare tests (bun run test:workers). Agent evals live inworkers/main/tests/evals/.e2e/— Playwright (bun run test:e2e).
- For UI route/component changes, run at least
bun run typecheckand the most relevant Vitest test(s). - For Worker/DO behavior, prefer focused
bun run test:workers -- <test-file>orbun run test:workerswhen the surface is shared. - For changes crossing browser chat, worker routing, and project runtime behavior, test the smallest representative path plus typecheck.
- Add tests when changing auth, billing/usage, admin purge/ban behavior, proxy auth, file safety, or persistence semantics.
sandbox/validate-notebook.pyis canonical;workers/main/analysis-sandbox-assets/validate-notebook.pymust stay byte-identical (tests/analysis-sandbox-asset-drift.test.ts).
Error Handling Culture
- Prefer failing loudly and early over silently swallowing errors or falling back to unclear behavior. Hidden failures make production bugs much harder to debug.
- Only add fallbacks when they preserve a clearly defined user experience and still expose enough signal through errors, logs, or tests to diagnose the original failure.
- Do not convert unexpected persistence, auth, upload, billing, or runtime/tool failures into empty data unless the caller explicitly treats "not found" as a valid state.
Local Environment Notes
Minimal prerequisites: Node.js 22+, Bun, Tailscale, and Cloudflare credentials for deployed/bound services.
Common local secret/config files:
.dev.varsfor Worker/dev secrets.wrangler*.jsoncfor environment-specific Cloudflare config. Useful local variables includeCF_GATEWAY_TOKEN, OAuth client IDs/secrets,INTEGRATION_SECRET_KEY,TOKEN_SIGNING_SECRET, and email provider settings.
SSH to shared hosts
Use Tailscale SSH as user chiridion; do not rely on shared private keys for normal access:
tailscale ssh chiridion@chiridion-sandbox-staging
tailscale ssh chiridion@chiridion-sandbox-prod
Tailscale host IPs are staging 100.115.221.105 and prod 100.112.135.2. Public SSH ingress should remain closed except temporary break-glass access.
Maintenance Rules
- Keep this guide short. Prefer pointers to files and tests over duplicating implementation details.
- When adding a major subsystem, add a short map entry and the canonical test command.
- When removing or renaming a subsystem, update this file in the same change.
- If a detail is likely to drift quickly, document where to verify it instead of freezing it here.
Cursor Cloud specific instructions
Durable, non-obvious notes for running this repo inside a Cursor Cloud VM (no Cloudflare account/credentials available). Standard commands live in the README.md / package.json tables above — this section only captures the gotchas.
- Run fully offline with
E2E_LOCAL=1. The defaultbun run devmarks several bindingsremote: true(AI,R2_BUCKET,ARTIFACTS,BROWSER,send_email) and expects a Cloudflare login. SettingE2E_LOCAL=1forces every binding into local Miniflare and disables sandbox containers (seevite.config.ts), so no Cloudflare creds and no Docker are needed for the web app. UseE2E_LOCAL=1 bun run dev:local-auth(auto-seeds aLocal Devuser/org/workspace and auto-logs-in; localhost-only) → app onhttp://localhost:3001. .dev.varsis required to boot and must contain non-emptyTOKEN_SIGNING_SECRETandINTEGRATION_SECRET_KEY(random values are fine — copy.dev.vars.example). It is gitignored; the setup step already created one in the snapshot, so it normally persists across sessions. Recreate it if missing.- Exercising a real chat turn offline (no model creds): run the deterministic fake LLM
node scripts/fake-llm.mjs(port8788) and start the dev server withTEST_LLM_REPLAY_URL=http://localhost:8788soresolvePiModelroutes model calls to it. The fake echoes text afterReply with:(prefixed[fake-llm]). Full hello-world command:TEST_LLM_REPLAY_URL=http://localhost:8788 E2E_LOCAL=1 bun run dev:local-auth - Local Playwright E2E:
E2E_LOCAL=1 bun run test:e2ereuses an already-running:3001dev server (and:8788fake LLM) viareuseExistingServer; if you want Playwright to own both, stop your manual servers first. Chromium + system deps are already installed in the snapshot (bunx playwright install --with-deps chromiumto refresh). The twoconnections-localspecs can flake against a live Vite dev server (dialog open timing); the app page itself loads fine. - CI does not gate
typecheckorlint(.github/workflows/ci.ymlruns onlytest:run,test:workers, and self-host checks). As a resultbun run typecheckandbun run lintcurrently report pre-existing failures onmain(e.g.tests/container-sizing.test.tsTS5097; ano-unused-varswarning inchat-thread-do.ts) that are unrelated to environment setup — do not treat them as regressions you introduced. - Bun is installed at
~/.bun/bin/bunand symlinked to/usr/local/bin/bunso it resolves in non-login shells (the update script'sbun install --frozen-lockfiledepends on this).
