Imported from kdlbs/kandev (
apps/backend/AGENTS.md). Install upstream withnpx skills add kdlbs/kandev --skill backend. Copyright stays with the author.
Backend (Go) — architecture and conventions
Scoped guidance for apps/backend/. Repo-wide rules (commit format, code-quality limits, etc.) live in the root AGENTS.md. For plugin work, start at the canonical plugin authoring guide, follow choose recipe → edit manifest.yaml → implement → validate → package → smoke test, and treat pkg/pluginsdk, proto/kandev/plugin/v1/plugin.proto, internal/plugins/manifest, and internal/plugins/pkgtar as authoritative; cmd/plugin-fixture is test support, including the remote executor provider and fake HTTPS service, and plugins must not access databases or internal/... packages.
Package Structure
apps/backend/
├── cmd/
│ ├── kandev/ # Main backend binary entry point
│ ├── agentctl/ # Agentctl binary (runs inside containers or standalone)
│ └── mock-agent/ # Mock agent for testing
├── internal/
│ ├── agent/
│ │ ├── runtime/ # Agent runtime: single seam for Launch/Resume/Stop/observe
│ │ │ ├── lifecycle/ # Agent instance management (moved from agent/lifecycle)
│ │ │ ├── agentctl/ # HTTP client for talking to agentctl (moved from agentctl/client)
│ │ │ └── routingerr/ # Provider error classifier + sanitizer + ProviderProber registry
│ │ ├── agents/ # Agent type implementations
│ │ ├── controller/ # Agent control operations
│ │ ├── credentials/ # Agent credential management
│ │ ├── discovery/ # Agent discovery
│ │ ├── docker/ # Docker-specific agent logic
│ │ ├── dto/ # Agent data transfer objects
│ │ ├── executor/ # Executor types, checks, and service
│ │ ├── handlers/ # Agent event handlers
│ │ ├── kubernetes/ # Kubernetes config, Pod/PVC composition, admission, and client streaming
│ │ ├── registry/ # Agent type registry and defaults
│ │ ├── settings/ # Agent settings
│ │ ├── mcpconfig/ # MCP server configuration
│ │ ├── remoteauth/ # Remote auth catalog and method IDs for remote executors/UI
│ │ └── planinjection/ # Bounds a task plan document before session-handover/dynamic-continuation injection
│ ├── auth/ # Opt-in auth, per-user scoping, middleware, API, store
│ ├── canvas/ # Agent-authored plugin web-app canvas lifecycle and governance
│ ├── agentctl/
│ │ └── server/ # agentctl HTTP server
│ │ ├── acp/ # ACP protocol implementation
│ │ ├── adapter/ # Protocol adapters + transport/ (ACP, Codex, OpenCode, Copilot, Amp)
│ │ ├── api/ # HTTP endpoints
│ │ ├── config/ # agentctl configuration
│ │ ├── instance/ # Multi-instance management
│ │ ├── mcp/ # MCP server integration
│ │ ├── process/ # Agent subprocess management
│ │ ├── shell/ # Shell session management
│ │ └── utility/ # agentctl utilities
│ ├── orchestrator/ # Task execution coordination
│ │ ├── dto/ # Orchestrator data transfer objects
│ │ ├── executor/ # Launches agents via lifecycle manager
│ │ ├── handlers/ # Orchestrator event handlers
│ │ ├── messagequeue/ # Message queue for agent prompts
│ │ ├── queue/ # Task queue
│ │ ├── scheduler/ # Task scheduling
│ │ └── watcher/ # Event handlers
│ ├── task/
│ │ ├── controller/ # Task HTTP/WS controllers
│ │ ├── dto/ # Task data transfer objects
│ │ ├── events/ # Task event types
│ │ ├── handlers/ # Task event handlers
│ │ ├── models/ # Task, Session, Executor, Message models
│ │ ├── repository/ # Database access (SQLite)
│ │ └── service/ # Task business logic
│ ├── runs/ # Generic run queue: models, repository, service, scheduler
│ ├── office/ # Autonomous agent management (agents, approvals, channels, config, configsync,
│ │ # costs, dashboard, infra, labels, onboarding, projects, repository, runtime,
│ │ # routines, routing, scheduler, service, shared, skills, workspaces)
│ ├── events/ # Event bus for internal pub/sub
│ ├── gateway/ # WebSocket gateway, including task-owned LSP lease lifecycle
│ ├── github/ # GitHub API integration (PRs, reviews, webhooks)
│ ├── githubauth/ # Shared GitHub credential-broker environment contract
│ ├── common/ # Shared utilities, config, logger
│ ├── integration/ # External integrations
│ ├── integrations/ # Shared shapes for third-party integrations
│ │ ├── healthpoll/ # Reusable 90s auth-health Poller (used by jira, linear)
│ │ └── secretadapter/ # Upsert-style adapter over secrets.SecretStore
│ ├── i18n/ # Localization for backend-rendered browser/share artifacts
│ ├── jira/ # Jira/Atlassian Cloud integration (config, REST client, poller)
│ ├── kubernetes/ # Task-owned Pod/PVC diagnostics; session stop preserves compute. See docs/specs/executors/system-design/kubernetes-task-pod.md for ownership, credentials, and recovery invariants.
│ ├── linear/ # Linear integration (config, GraphQL client, poller)
│ ├── lsp/ # LSP server
│ ├── mcp/ # MCP protocol support
│ ├── health/ # Health check endpoints
│ ├── notifications/ # Notification system
│ ├── persistence/ # Persistence layer
│ ├── prompts/ # Prompt management
│ ├── repoclone/ # Repository cloning for remote executors
│ ├── scriptengine/ # Script placeholder resolution and interpolation
│ ├── secrets/ # Secret management
│ ├── sprites/ # Sprites AI integration
│ ├── sysprompt/ # System prompt injection
│ ├── tools/ # Tool integrations
│ ├── user/ # User management
│ ├── utility/ # Shared utility functions
│ ├── workflow/ # Workflow engine (engine, models, repository, service)
│ ├── workflowsync/ # GitHub workflow sync (per-workspace repo config, poller, force sync)
│ └── worktree/ # Git worktree management for workspace isolation
Canvas creation authority comes only from the trusted task adapter and binds
the owner, session, task, and policy version. Its first valid static release
may consume that authority for exact workspace-ceiling grants in the activation
transaction. Existing drafts, imports, later permission increases, and
revoked grants still require human review; source metadata and manifests never
grant authority.
Canvas scopes expose canvas_data_scope_transition_total with fixed transition
and result labels. Lifecycle events carry IDs; metrics never label task data.
Key Concepts
Orchestrator coordinates task execution:
- Receives task start/stop/resume requests via WebSocket
- Delegates to lifecycle manager for agent operations
- Handles event-driven state transitions via workflow engine
- Located in
internal/orchestrator/
Cancellation progress projection: orchestrator.Service.CancellationPending(sessionID) is a runtime-only, session-scoped view of accepted cancellation work. Serialization that carries the boolean with ordering identity uses the atomic CancellationPendingSnapshot(sessionID) provider, whose process-local revision increments on first-begin and last-end transitions.
The task DTO package exposes both the compatibility boolean provider and snapshot seam; boot state, task-session HTTP/WS lists and detail responses, and the session-scoped WebSocket notification must project explicit true/false values plus the revision.
Keep count, revision, and publication queue updates in one critical section, drain event-bus sends outside it, and never persist this transient marker or turn it into a coarse session lifecycle state.
Watcher Dispatch Coordinator (internal/orchestrator/watcher_dispatch.go) is the single pipeline that turns a freshly-observed external issue (Linear, Jira, future) into a Kandev task. Bus subscribers for each integration forward the event to WatcherDispatchCoordinator.Dispatch with a per-integration WatcherSource implementation (source_linear.go, source_jira.go). Source methods carry the integration-specific bits (reserve dedup, build task request, attach task ID, release, auto-start params); the coordinator owns the cross-cutting pipeline (create task, decide auto-start, error/release handling). Add a new watcher = implement WatcherSource + register a one-line bus subscriber. Do NOT add another createXIssueTask mirror.
GitHub App registration catalog (internal/github/) stores zero or more managed/imported App
registrations; none is a global default. Workspace App connections must carry both registration ID
and installation ID. Runtime clients, token caches, broker leases, OAuth state, webhook delivery,
and service cache scopes must retain registration identity and credential generation. Reusing a
registration intentionally shares root App credentials and bot identity; installation grants and
workspace credentials remain isolated. Registration create/import/select/install belongs to the
workspace GitHub settings flow, and backend startup configuration is not an App credential source.
Registration-specific callback and webhook routes select one candidate registration but never
replace state verification, installation association, or HMAC verification.
Workflow Engine (internal/workflow/engine/) provides typed state-machine evaluation:
Engine.HandleTrigger()evaluates step actions for triggers (on_enter, on_turn_start, on_turn_complete, on_exit)TransitionStoreinterface abstracts persistence (implemented byorchestrator.workflowStore)CallbackRegistrymaps action kinds to callbacks (plan mode, auto-start, context reset)- First-transition-wins: multiple transition actions in one trigger, first eligible wins
EvaluateOnlymode: engine evaluates without persisting, caller orchestrates on_exit → DB → on_enterRequiresApprovalon actions: transitions requiring review gating are skipped- Idempotent by
OperationID; session-scoped data bag viaMachineState.Data
Agent Runtime (internal/agent/runtime/) is the single seam for launching, resuming, stopping, and observing agent executions. ADR 0004 introduced this in Phase 1 of task-model-unification. The public surface is runtime.Runtime (runtime.go); a thin facade (facade.go) delegates to a Backend (satisfied by *lifecycle.Manager). Run-owned executions use runtime.LaunchSpec.Owner (kind=run) with durable run-session identity; admission fails closed before allocation and lifecycle registration, and runtime.Start rolls back failed startup while task launches keep task/session checks.
Run scheduling ownership: internal/runs/ is generic; internal/runs/models owns the shared run-row and run-event data contracts. Only internal/backendapp/ constructs and owns the single internal/runs/scheduler and its lifecycle. Office adapters may depend on runs, but generic runs must not import internal/office or its subpackages. Office retains launch, causation, routing, and backpressure policy. See ADR 2026-09-26-run-contract-ownership.
Runtime environment invariant: Agent.Runtime().Env applies to every ACP subprocess entry point. Route new overrides through host-utility probes and sessionless prompts into agentctl child processes before sanitization; cover probe DTO, prompt DTO, and child-process boundaries. Host utility probes must use the same profile-resolved HOME, GH_CONFIG_DIR, and credential-selection inputs that the eventual launch receives; compose structured env blocks once and propagate them across create/configure/start/reconfigure, test both profile/host mismatch directions, and never emit secret or token values.
System storage cleanup: Install-wide cleanup providers live under internal/system/storage/ and share the storage mutation gate for conflicting cache operations. Go-cache cleanup deletes eligible build data in place, preserving its root, ownership marker, and root fuzz corpus. It validates mount identity at traversal and removal boundaries; a missing mount identity must fail closed. Threshold discovery and deletion are bounded and resume across calls in the same process. The persisted go_cache.allow_cleanup_while_busy option bypasses activity and idle admission only for Go-cache cleanup; other providers retain their existing gates.
Convention: only internal/agent/runtime/ (and code that pre-dates Phase 1 migration) may import runtime/lifecycle or runtime/agentctl directly. New consumers — workflow engine actions, cron-driven trigger handlers, future task-tier callers — should depend on runtime.Runtime or narrow local interfaces for the lifecycle-owned capability they consume. Existing call sites are migrated through later phases of task-model-unification.
Lifecycle Manager (internal/agent/runtime/lifecycle/) manages agent instances under the runtime:
Manager(manager.go,manager_*.go) - central coordinator for agent lifecycleExecutorBackendinterface (executor_backend.go) - abstracts execution environment (Docker, Standalone, Sprites, SSH, Kubernetes, Remote Docker)ExecutionStore(execution_store.go) - thread-safe in-memory execution trackingsession.go- ACP session initialization and resumestreams.go- WebSocket stream connections to agentctlprocess_runner.go- agent process launch and managementprofile_resolver.go- resolves agent profiles/settings Lifecycle callback identity: Process callbacks must retain the launched PID and generation captured when scheduled, revalidate both before state/I/O, and ignore delayed callbacks from replaced processes and duplicates; test replacement before start and after waits. Prompt callbacks and asynchronous prompt errors must likewise carry immutable execution, session, prompt-generation, and turn evidence captured at the result boundary; never reread mutable execution snapshots later, because a successor prompt may already own them. Settle through the correlated terminal path and test both replacement-execution and same-execution successor-prompt races. Event handlers must not synchronously call lifecycle-manager methods that can reacquire a lock held by the event publisher; terminal event payloads must carry immutable identity/admission evidence for that boundary. Add a regression that blocks any manager generation read while publishing the terminal event.
agentctl client (internal/agent/runtime/agentctl/) is the HTTP/WS client used by the lifecycle manager to talk to a running agentctl instance. It is a runtime-tier package and should not be imported outside internal/agent/runtime/.
Agent discovery vs. ACP probing: discovery answers only whether an agent executable is available; authentication, protocol compatibility, and supported models or modes belong to the ACP probe path, not installation gates. MiniMax uses native mcode acp, dynamic encoded model IDs and executor-owned ~/.minimax; do not relocate its auth files because their identity includes the absolute auth-home path.
agentctl is an HTTP server that:
- Runs inside Docker containers or as standalone process
- Manages agent subprocess via stdin/stdout (ACP protocol)
- Exposes workspace operations (shell, git, files)
- Supports multiple concurrent instances on different ports
Standalone agentctl is launched in its own process group so terminal Ctrl+C is handled by the backend lifecycle manager first; do not share the backend's foreground group, which bypasses supervised shutdown and can leak ACP subprocesses.
Executor Types (database model):
local_pc- Standalone process on hostlocal_docker- Docker container on hostsprites- Sprites cloud environmentssh- Remote SSH hostk8s- Namespaced Kubernetes Pod with optional PVC workspaceremote_docker- Container on a Docker daemon reached over SSH;remote_vps- Planned
Kubernetes lifecycle: task_environment_kubernetes owns shared physical Pod/PVC inventory; executors_running records individual sessions and legacy session-owned pods. Persist the exact Pod/PVC names, UIDs, full kandev.ai/* identity, workload snapshot, and internal runtime-secret references before reporting a launch as durable. Session stop deletes only its agentctl instance, and backend shutdown preserves resources; task cleanup deletes the Pod and only a Kandev-created PVC after exact identity checks and confirmed absence. Reconnect uses the current executor connection config but the recorded workload/resource snapshot, and any ambiguity fails closed. Keep agentctl reachable only through a process-local loopback port-forward; never add a Service or place resolved credentials in a Pod spec.
Remote SSH executor platforms: Treat supported remote OS/arch values as an end-to-end contract. Platform probe/normalization, lifecycle support checks, agentctl helper resolution, platform default shell, SSH readiness endpoints, frontend response types, and tests must stay aligned. Preserve raw unsupported platform details in user-facing errors, but use normalized values for supported-platform matching. Keep shell defaults platform-aware: Darwin defaults to zsh, Linux defaults to bash, unless an explicit shell is saved.
Remote SSH lifecycle: Resolve and run remote preparation, then verify the canonical checkout, origin, and HEAD before starting agentctl. Retain the remote task directory in session state for stop/resume. On a terminal archive/delete/cascade stop the profile's cleanup_script runs over the live SSH connection before teardown; StopInstance never removes the task directory itself, on any stop reason. Ordinary stop and backend shutdown preserve the workspace. Reclaiming the task directory is an opt-in phase of the durable task-resource cleanup job, not part of the stop path (see docs/specs/executors/system-design/remote-task-directory-reclamation.md). Every stop reason except graceful backend shutdown kills the remote agentctl process and removes its per-session runtime directory.
Opt-in authentication & per-user scoping (internal/auth/): auth is OFF by default; the global middleware (auth/httpmw, installed after CORS in backendapp.buildHTTPServer) then injects a synthetic admin identity for the pre-auth single user, so behavior is unchanged. Enablement is the features.auth runtime flag (KANDEV_FEATURES_AUTH, Settings > System > Feature Toggles) — the auth service derives its mode from cfg.Features.Auth (disabled / setup = flag-on-no-admin / enabled = flag-on-admin-exists); there is no separate auth.mode setting. When enabled, requests authenticate via the session cookie (base name kandev_session; the effective name is derived from the request host — port-scoped on a ported host via internal/common/httpcookie, plain on a default-port host) or a kandev_pat_* bearer token. Scoping rules:
- Identity travels in the request context (
authn.IdentityFromContext). No identity = internal caller (pollers, event bus, office schedulers) = unscoped. Synthetic identity = auth disabled = unscoped. - In-session agent MCP is scoped to the task owner. The MCP tools an agent gets inside its own session are relayed over the agent's WebSocket stream, which carries no credential of its own.
internal/mcp/scoperesolves the stream's task → workspace → owner and attaches that user's real identity before dispatch (lifecycle.Manager.SetMCPIdentityScoper), so the sameauthorize*checks apply as for the PAT-authenticated/mcpendpoint. The owning task comes from theAgentExecution, never from the agent-supplied payload — do not "improve" this by readingsession_id/task_idout of the request. Tool handlers stay identity-agnostic. Under enforced auth every dispatch is scoped to somebody: a resolvable active owner gets their identity; an unowned workspace gets a sentinel user ID that reaches unowned rows only; anything unresolvable (missing task/workspace row, or an owner whose account was deleted or disabled) is denied. Never return an identity-free context from this path — the task service reads that as an internal caller and grants everything. - Workspaces are per-user (
workspaces.owner_id); the task service filters/denies at the service layer with*NotFoundsentinels (no existence leak). Unowned rows (owner_id='') stay visible until the setup wizard claims them for the admin. - New user-facing service entry points must apply scoping — call the
authorize*helpers intask/service/service_access.go(or the same pattern) when adding routes that read or mutate workspace-scoped data. This includes session-keyed entry points: any service method taking a caller-suppliedsessionID(or ataskIDused to reach sessions) authorizes viaauthorizeTaskID(session.TaskID). Authorize off the row you already read rather than callingAuthorizeSessionAccess, which re-reads the session. Batch helpers that take ID lists (BatchGetSessionsForTasks,GetPrimarySessionInfoForTasks, …) are exempt by convention: their callers derive the IDs from an already-authorized task list, so per-ID checks would add N queries to list views for no gain — keep it that way, and don't hand them caller-supplied IDs. - Two session paths bypass the task service and must guard themselves. (1) Handlers that resolve an execution by a bare in-memory lookup (
GetExecutionBySessionID,*BySessionID) skip theGetOrEnsure*chokepoint where the lifecycle check runs — callservice.AuthorizeSessionAccess(seeProcessHandlers.denySessionAccess) orlifecycleMgr.CheckSessionAccessfirst. (2) The orchestrator reads sessions through its own repo handle, so it inherits nothing from the task service; session-keyed entry points there callauthorizeSession/authorizeTask, wired bySetSessionAccessChecker/SetTaskAccessChecker. An entry point that accepts both a task and a session ID must useauthorizeTaskSessionPair, neverauthorizeSessionalone: the caller can satisfy a session-only check with one of their own sessions while pointingtaskIDat someone else's task, which the method then uses for its task-scoped work. Put the guard first in the method, before any repo/executor/agent-manager use —TestSessionKeyedEntryPointsGuardBeforeDependenciesruns each entry point with nil dependencies and fails on a panic if a guard is placed too late. - The workflow-step surface scopes itself in
internal/workflow. Steps, workspace step lists, export and import all reach a workflow or workspace by caller-supplied ID, and none of those IDs is a name the WS dispatch backstop parses, so the guard has to be in this package.internal/workflow/service/access.goholdsAuthorizeWorkflow/AuthorizeWorkspace/AuthorizeStep, backed bySetWorkflowAccessChecker/SetWorkspaceAccessChecker(wired totaskSvc.AuthorizeWorkflowAccess/AuthorizeWorkspaceAccessinbackendapp/services.go, next to the olderSetSessionAccessChecker). Step CRUD authorizes in the controller, because REST, WS and the MCP tools all share it;ListStepsByWorkspaceIDand the export/import trio authorize in the service, because the boot-payload builder and the MCP config tools call those directly. A denial and a genuine miss both surface asservice.ErrNotVisible→ one 404, so a foreign step and a nonexistent one are indistinguishable. Templates are deliberately unscoped (install-global, ownerless). Note that an empty ID fails closed here rather than being passed toAuthorizeWorkspaceAccess, which reads""as "no scoping applies". A step also contains IDs —move_to_stepnames a step,queue_runnames a task — and the engine dereferences them later on the event bus, with no identity to check against, so they are validated at write time:models.CollectStepEventReferencesenumerates them (keep it in sync withRemapStepEventswhen adding a trigger), a transition target must resolve to a step in the same workflow, and a task target is authorized like any other. A reorder's step IDs must belong to the workflow named in the URL, which is also what keeps a read-only workflow's steps from being reordered through a mutable workflow's route. - Dispatched WS actions have a gateway backstop.
Client.authorizeAction(internal/gateway/websocket/dispatch_scope.go) checks any payload carryingtask_idorsession_idagainst the client's identity before dispatch, so a newly added action is scoped by default rather than only when its author remembers. Handler-level scoping is still required (defense in depth, and it covers non-WS callers) — do not delete a service-layer check because the backstop exists. The backstop readstask_id,session_idandtask_environment_id, plusidon top-leveltask.<verb>actions only (task.state,task.move,task.get, … —task.and exactly one more segment). Keep using those names for task-scoped actions, because an action that invents a new name for the same kind of resource silently opts out. (user_shell.stopwas exactly that: it keys offtask_environment_idwithtask_idoptional, so atask_id-only backstop missed it.task.stateandtask.movewere the second instance: they name the taskid, so the backstop parsed no refs and let one user mutate any other user's task — hence theidrule.) Deepertask.*namespaces are deliberately excluded: intask.plan.revision.getandtask.review.finding.updatetheidnames a revision or a finding, so those actions must carrytask_idwhen they are task-scoped. Also avoid permissive types (any,json.RawMessage) for those fields: a payload that fails to unmarshal into the ref struct is treated as naming nothing. The backstop does not cover HTTP. The dispatch backstop only sees WS payloads, so an HTTP route that names a task or environment has to guard itself, and only the third-party integration prefixes have a global middleware. The two SSR terminal lists (GET /api/v1/environments/:id/terminals,GET /api/v1/tasks/:id/terminals) were the gap: they read straight from the interactive runner and the terminal service (which has no authorization of its own), leaking another user's terminal display names and initial command lines. They now mountauthorizeEnvironmentRoute/authorizeTaskRoute(internal/agent/handlers/shell_handlers.go) as gin route guards callinglifecycleMgr.CheckEnvironmentAccess/CheckTaskAccess, so the check runs before any terminal state is read. Guard every ID the handler consumes, not just the path param:/tasks/:id/terminalsalso takes?task_environment_id=, which is a second way into a foreign environment. When a handler merges state keyed by two IDs, authorize them as a pair (AuthorizeTaskEnvironmentAccess, the task/environment sibling ofAuthorizeTaskSessionAccess) rather than independently: both single-ID checks pass for a caller who holds each ID separately, and the handler then merges two unrelated lists. Pair on the binding, not on row ownership:inherit_parentbinds a subtask's session to the parent task's environment andshared_groupbinds every group member to one canonical environment, soenv.TaskID == taskIDis the wrong predicate and would empty the terminal panel for both modes. - WS: clients carry their identity; dispatched actions and subscriptions are scoped; workspace-carrying events route via
Hub.BroadcastToWorkspace. A newhub.Broadcast(global) call site needs a//ws:globaljustification comment. Subscription actions (*.subscribe)returnbefore reaching the dispatch backstop (Client.authorizeAction), so scoping a new subscribe topic means adding a hook toSubscriptionAccessPolicy(internal/gateway/websocket/access.go) and calling it explicitly inside that action's own handler — the backstop'sscopedActionRefscannot reach it.run.subscribe'sSubscriptions.Runhook (added for WO-02, resolving the run's workspace viaruns.GetRunWorkspaceID) is the worked example. - Task overview vs. session detail: Keep cross-task state in the bounded, persisted
TaskStatusSummaryprojection andtask.status_summary.updated; keep unbounded transcript and execution data session-scoped. The bounded task-status spec and accepted ADR define the shared contract. - Port-proxy subtree capability: with auth enforced,
/port-proxy/:sessionId/:port/*needs a credential on every request, but browser subresource fetches cannot always carry one (the<link rel="manifest">fetch never sends cookies; sandboxed iframes and?token=-opened previews drop them elsewhere).PortProxyHandlermints a short-lived (15 min, sliding), path-scoped, HMAC-signed capability after the document authenticates and propagates it as a cookie scoped to the proxy subtree plus akandev_capquery parameter appended to rewritten asset URLs.requireConnectionAuth(internal/gateway/websocket/access.go) accepts either form only for the exact session:port it was minted for and restores the issuing identity, soCheckSessionAccessstill gates on the real owner. Synthetic identities (auth disabled) skip minting — proxied bodies stay byte-identical to the pre-auth behavior. - Self-authenticating webhooks (automation and office channels) and
/health,/ready,/api/v1/features,/api/v1/app-statestay public — the allowlist lives inauth/httpmw/middleware.gowith a pinning test. Plugin webhooks are not allowlisted:auth/httpmwstructurally defers only GET/POST relay paths because it cannot read manifests, andinternal/plugins.Controller.webhookenforces each manifest'saccesspolicy.
Execution Flow
Client (WS) → Orchestrator → Lifecycle Manager → ExecutorBackend (container/process) → agentctl
↓
Client (WS) ← Orchestrator ← Lifecycle Manager ←──── stream updates (WS) ──────── agent subprocess
- Orchestrator receives
session.launchvia WS - Lifecycle Manager creates executor instance (container or process)
- agentctl starts inside the instance, agent subprocess is configured and started
- Agent events stream back via WS through the chain
Session Resume: TaskSession.ACPSessionID stored for resume; ExecutorRunning tracks active state; on restart RecoverInstances() reconnects.
Provider Pattern: Packages expose Provide(cfg, log) (*impl, cleanup, error) for DI. Returns implementation, cleanup function, and error. Cleanup called during graceful shutdown.
Worktrees: internal/worktree/Manager provides workspace isolation. Each session can have its own worktree (branch) to prevent conflicts between concurrent agents.
Worktree file materialization: copy_files (Repository.CopyFiles) is a comma-separated spec of repository-relative gitignored paths/doublestar globs seeded into each new worktree. Copy is the default; the exact terminal :symlink suffix (for example .env.local:symlink) creates a relative host-worktree link so source changes propagate live. Other colons stay literal (config:dev, .env:); use ::symlink to copy a literal path ending in :symlink. Malformed reserved syntax is rejected at save time. Parsing and materialization live in internal/worktree/copyfiles/ (ParseSpecs, ValidateSpec, Copy); duplicates and overlapping matches are first-entry-wins. Manager.copyConfiguredFiles runs before setup during worktree creation. Source and destination containment checks reject traversal and symlinked destination parents. Failures are non-fatal warnings. Windows link creation is best-effort. Remote executors cannot link to the host, so Parse/Plan preserve literal colon paths but turn symlink-mode entries into copied bytes delivered through WriteEntries.
Executor default scripts: Default prepare scripts are in internal/agent/runtime/lifecycle/default_scripts.go; internal/scriptengine/ handles placeholder resolution.
- Launcher quoting: Quote interpolated environment values for POSIX/Git Bash
versus native Windows; when launch recipes change, extend
scripts/check-make-shellswith default, spaced, and quote-containing values.
Conventions
- Provider pattern for DI; stderr for logs, stdout for ACP only.
- Pass context through chains; event bus for cross-component comm.
- Dependency direction: Shared admission limits used by task models and orchestrator/messagequeue belong in a neutral internal package; task/models must not import the higher-level orchestrator/messagequeue package.
- Production Git commands must use
subproc.NewGitCommandwith a classifiedsubproc.RunGit*helper, or hold a classified admission slot across streamingStart/Wait. Do not construct raw Git commands outsideinternal/common/subproc; chooseinteractive,lifecycle, orbackground. - Managed Git final preparation runs immediately before
Startand owns the fullStart-to-Waitlifecycle. Network callers choose finite post-admission budgets; useAfterAcquirebuilders so queue wait is outside the execution budget. Shared lifecycle code belongs ininternal/common, and agentctl keeps only compatibility wrappers. - Cross-tier shared code belongs in
internal/common/. agentctl ships as a standalone binary uploaded into containers, so neither it nor the backend may own code the other imports — put shared logic in a neutralinternal/common/*package instead of importing across the boundary or forking a copy (ptyexecis the worked example;subproc,gitref,securityutilare older ones). Duplicating to "avoid the dependency" is how the two PTY copies drifted. - Startup configuration ownership:
internal/common/config/catalog.goandsource.goown the stable operator catalog, discovery, precedence, provenance, and typed values; managed agentctl children receive a private subset contract. Do not reparse stable env in executors or copy YAML into public child environments. - Pure computation that only matters on Windows should not carry
//go:build windows. One CI job runs on Windows and it tests a package allowlist, so tagged code is easily unverified — loginpty's copy of the Win32 quoting was compiled by no job at all. Keep platform API calls behind the tag, leave string/path helpers untagged, and add packages needing native coverage to thetest-windowsjob in.github/workflows/backend-tests.yml. - Event-bus wildcard parity: New NATS wildcard subscriptions must verify equivalent
MemoryEventBussemantics ingo test ./internal/events/bus.EventBus.Publishcompletion means transport/enqueue, not remote NATS subscriber processing; never use publication as a local ordering acknowledgement. If a waiter must not release until local state is armed, use a construction-supplied synchronous local notifier or durable state transition before publishing fan-out, and test with an async/NATS-like bus fake as well asMemoryEventBus. - Repository provider identity: Provider-backed repositories are keyed by workspace, provider, normalized
provider_hostorigin, full owner/namespace, and name. Persistprovider_hostwhen importing or resolving a remote; do not infer self-managed GitLab rows from owner/name alone. Legacy rows with an empty host have unknown identity and must fail closed for provider write/link operations. - Repository provider branches: Native task branch pickers derive provider identity only from the persisted workspace repository. GitHub uses the first-party service; manifest-owned providers route through the owning active plugin's workspace-scoped
repositories.branchesaction. Do not add provider-specific task-service branches or accept browser-supplied repository descriptors on this path. - PR status sync: For watch identity, PR lifecycle persistence, or batched lookup changes, read GitHub guidance.
- Execution access: Workspace-oriented handlers (files, shell, inference, ports, vscode, LSP) MUST use
GetOrEnsureExecution(ctx, sessionID)— it recovers from backend restarts by creating executions on-demand. Only useGetExecutionBySessionIDfor operations that require a running agent process (prompt, cancel, mode). - Generated-title ownership: Claim generated-title ownership only after task, config, and Office eligibility is known. Config tasks, Office/External modes, and workspace-only preparation (
StartAgent: false) must never claim it; only an eligible agent-start task may claim the generated title. - Task writes: Canonical parent admission uses the transaction-bound reader in
internal/task/repository/hierarchy. Reserve SQLite's writer before graph reads; on PostgreSQL use READ COMMITTED and lock sorted workspaces before sorted steps and task rows. Full task snapshots preserve the current parent and normalized materialized workspace mode/group; only admitted explicit parent intent changes the edge. Office scalar parent updates participate in serialization while retaining their deliberate deeper-cycle/depth permissiveness. OrdinaryService.UpdateTaskrequests use requiredUpdateTaskFieldsWithParentAdmissionto overlay supplied fields on the locked current task, preserving omissions and metadata/title ownership. Legacy full-snapshot, exact and workflow writes retain their own contracts. - Task lifecycle events: Any code path that mutates a task row must publish via the event bus (
task.created/task.updated/task.deleted) — either by going throughService.CreateTask/UpdateTask/DeleteTask/ArchiveTask, or by callingpublishTaskEvent(or one of thePublish*helpers inservice_events.go) directly. Walkingrepository.TaskRepositorystraight bypasses event publishing and breaks WS-driven UI like the All-Workflows kanban view.HandoffService's cascade methods learned this the hard way — they now require aTaskEventPublisherwired viaSetTaskEventPublisher. New cascade / bulk / cleanup paths must follow the same pattern. Workflow steps follow the same rule with their own publisher: step create/update/delete publishworkflow_step.created/.updated/.deletedthroughinternal/workflow/stepevents.Publisher(the WS gateway fans them out; the orchestrator re-evaluates queue admission off.updated). Both mutation surfaces — REST/WS ininternal/workflow/handlers, MCP ininternal/mcp/handlers— share that publisher and differ only in the source label they pass, so a new step-mutation path uses it rather than hand-rolling the payload; promoting a start step demotes the previous one, so also publish.updatedfor every entry inDemotedStartSteps. - Agent task plans: Agent whole-document writes must carry
expected_versionwhen a plan already exists. The task service rejects suspicious reductions before storage and requires explicitallow_truncationplus verified history for intentional reductions. Exact edits and revision recovery stay task-scoped, use the service lock, and check both the current plan version and the selected revision snapshot. Keep browser DTOs and browser write behavior independent from the agent-only version contract. - Workspace deletion side tables: Any workspace-scoped integration side table without a database foreign key/cascade must be deleted explicitly from its
WorkspaceDeletedhandler. Cover the create → delete lifecycle in a test, and include the same cleanup in E2E reset fixtures; do not assume deleting the workspace row removes orphaned integration metadata. - Testing: For backend test changes, load backend-tests.md. It covers time, cleanup, filesystem fixtures, environment isolation, and subprocess helpers.
Goroutine ownership and leak testing
Every long-running goroutine must have a single owner with explicit start and stop semantics:
- Lifecycle: the type that spawns the goroutine also exposes
Start(ctx)/Stop()(or equivalent).Startregisters on async.WaitGroup;Stopcancels the goroutine's context (or closes astopCh) andwg.Wait()s for drain. Idempotent on both ends. Restartable services must reset stopped state and create a fresh cancellable context/cancel under the same mutex used byStop; coverStop→Start.internal/integrations/healthpoll,internal/jira,internal/linear, andinternal/githubpollers are the canonical shape. - E2E reset invariant:
seedData/backend are worker-scoped, so any workspace-scoped state a global poller reads (for examplegithub_review_watches) must be deleted incmd/kandev/e2e_reset.gobefore task deletion — otherwise the poller recreates rows mid-reset and later tests see duplicates. Add aDelete...ByWorkspacecascade when introducing a new poller-backed entity. - Cancellation: the goroutine selects on
ctx.Done()(orstopCh) in every long wait. Never usetime.Sleepin a retry/backoff loop — usetime.NewTimerinside aselectthat also watches the shutdown signal (seelifecycle.StreamManager.sleepOrStop). - Detached helpers: event handlers and short-lived
go func()calls ininternal/orchestrator/andinternal/agent/runtime/lifecycle/must accept a cancellable context (or check the owning type's shutdown signal) and return promptly when it fires. - Leak testing: packages that spawn goroutines add
goleak.VerifyTestMain(m)in a per-packageTestMain. New packages of this kind must follow suit. Tests that spawn polling goroutines must bound them with a context or timer and stop them viat.Cleanup. When a third-party background goroutine genuinely can't be drained, suppress it withgoleak.IgnoreTopFunction(...)and leave a comment explaining why. Currently instrumented:internal/gateway/websocket/,internal/agent/runtime/lifecycle/,internal/agentctl/server/process/,internal/orchestrator/,internal/github/,internal/gitlab/,internal/jira/,internal/linear/,internal/integrations/healthpoll/.
Backups
internal/system/toolretentionowns opt-in payload cleanup; policy and progress share one settings record. Its batches shareinternal/system/maintenanceadmission with backup, restore/reset, and compaction. Keep preparation cancellable and accepted manual jobs independent of HTTP cancellation. Message replacements must preserve removal markers and activity timestamps; analysis and cleanup share the reducer. Status polling and startup never scan payloads. See the retention design.- On every SQLite boot,
persistence.Providereadskandev_meta.kandev_version. If the stored version differs from the binary version (or any user tables exist but no version is recorded), it takes aVACUUM INTOsnapshot intobackups/beside the configured SQLite database file before running migrations.VACUUM INTOrecreates its target, so create staged snapshots inside a same-filesystem private0700staging directory and chmod the file to0600before validation/install; removing a placeholder and chmodding afterward does not protect the creation window. The default path is<home>/data/kandev.db, with backups in<home>/data/backups/. - Retention: 2 backups kept (newest two by mtime); older ones are pruned after the snapshot succeeds.
- Postgres: backup is skipped with a log line. Use
pg_dumpfor Postgres backups. - Boot aborts if the backup fails — the pool is closed and
Providereturns an error. - Restore quiesces scheduling, active executions, and database-backed workers, validates the SQLite checkpoint result, and closes the shared pool. It quarantines the configured database and sidecars, installs the staged file, and restores the originals if replacement fails. The frontend requires a restart before database-backed work resumes; PostgreSQL restore is rejected.
- After all repos complete
initSchema,cmd/kandev/storage.go:recordSchemaVersionwrites the current binary version intokandev_meta(non-fatal; a failure just means the next boot will take a fresh snapshot). - Migration logging:
db.MigrateLogger.Apply(name, stmt)— success logs Info, "already exists" / "duplicate column name" is silently swallowed, anything else logs Warn but never returns an error (preserving the existing swallow-error contract). - Schema replay handling: use
internal/dbhelpers such asIsDuplicateColumnError/IsAlreadyExistsErrorinstead of local error-string matching. When adding or changing startup schema code, include fresh-DB plus same-DB replay tests for SQLite; add the same env-gated Postgres replay coverage when the path supports Postgres. Seedocs/decisions/0027-replayable-schema-migrations.md.
Schema & migrations (SQLite repository)
internal/task/repository/sqlite also runs against PostgreSQL: avoid unguarded SQLite-only rowid, JSON, or date syntax and add an environment-gated PostgreSQL behavior test for every changed dialect-sensitive method (schema replay is insufficient). For PostgreSQL read-modify-write invariants, acquire a transaction-scoped pg_advisory_xact_lock(hashtextextended(namespace+identity, 0)) before the first read or update; SQLite's single-writer lock does not need this, and unique constraints/retries remain defense in depth. Verify with an env-gated real multi-connection test under go test -race.
initSchema() in internal/task/repository/sqlite/base_schema.go runs the init*Schema (CREATE TABLE) steps before runMigrations(). The table-creation DDL uses CREATE TABLE IF NOT EXISTS, so on an existing database it is a no-op and never adds columns to a table that is already present.
Rule: when you add a column to an existing table, add it only via an idempotent ADD COLUMN migration in runMigrations() (base_migrations.go), never by editing the table's CREATE TABLE alone. Anything that references that new column — an index, a backfill UPDATE, a partial-index predicate — must live in runMigrations() after the ADD COLUMN, not in the init*Schema DDL. Putting a CREATE INDEX ... (new_col) in the schema-init block crashes existing DBs with no such column: new_col, because schema init runs before the migration that adds the column.
You may still list the column in the CREATE TABLE so fresh DBs get it inline, but the migration is the source of truth for evolution and must stand alone. New columns also need: the struct field in models/, the DTO field + ToAPI in pkg/api/v1/, and every CreateX/UpdateX/bulk write in the repo that should set it.
Built-in prompt content refreshes are seed-data migrations, not schema migrations. Match only known historical content hashes after applying the same normalization as the embedded prompt loader, require created_at == updated_at to preserve user edits, and use a conditional update over the original row values to avoid racing concurrent edits. Keep these refreshes with prompt seeding rather than runMigrations().
Internationalization
internal/i18n renders only browser-facing copy: SPA-unavailable pages and shared-task artifacts. Diagnostics, logs, agent/ACP output, and CLI output remain English. Use i18n.T/i18n.Tf with explicit locale threading (including interpolation/plurals); resolve artifact locale at creation. Catalogs are embedded in internal/i18n/locales/; regenerate pseudo with pnpm run i18n:pseudo.
Prefer stable error codes for new output so the frontend translates it. See docs/i18n.md and ADR 2026-08-01-share-artifact-locale.md.
Table-rebuild migrations: When a legacy or constraint migration recreates a table, mirror every new column in the replacement CREATE TABLE and INSERT ... SELECT copy list; add a replay regression test proving values, including timestamps, survive. Destructive cutover migrations: Build a legacy schema with NewWithDB on a fresh database, replace final-shaped tables with legacy-shaped ones, inject a test-only failpoint after each cutover step, and assert byte-equivalent rollback; run the same matrix with KANDEV_TEST_POSTGRES_DSN, looking up PostgreSQL constraint names dynamically because they truncate at 63 bytes.
Every built-in SQL schema owner needs a descriptor in internal/persistence/requiredstores and a fixed adapter in internal/persistence/storeconformance. Bootstrap records each constructor through the tracker; missing required schema fails before readiness. Provider credentials and remote probes remain independently degradable.
For persistence changes, run go run ./cmd/sqlguard ./internal and go test -race ./internal/persistence/storeconformance -count=1; set KANDEV_TEST_POSTGRES_DSN for PostgreSQL coverage and update the explicit-tag upgrade fixture and manifest when schema history changes.
Code-quality limits
Enforced by apps/backend/.golangci.yml (errors on new code only):
- Functions: ≤80 lines, ≤50 statements · Cyclomatic complexity: ≤15 · Cognitive complexity: ≤30 · Nesting depth: ≤5 · Naked returns only in functions ≤30 lines · No duplicated blocks (≥150 tokens) · Repeated strings → constants (≥3 occurrences) · Revive's 800-effective-line file limit also applies to test files; put new tests in a new file instead of appending to an already-large test file.
When a PR fixup touches backend code, run golangci-lint run ./... --new-from-rev="<base-sha>" --timeout=5m from apps/backend with the PR base SHA before pushing; CI enforces changed-file complexity thresholds.
- Build cache reuse: Make build/test targets use
-trimpath. Include it in directgo buildandgo testcommands, including race/coverage runs, so identical packages share artifacts across worktrees. Diagnostic source paths use module paths instead of absolute worktree paths. internal/launcher/— native launcher owning every entrypoint (dev,start,run,service);devrunsmake -C apps/backend devwith Vite as a supervised child, state under<repoRoot>/.kandev-dev/. The rootmake devprebuilds only the copied launcher; the backend dev target builds the native agentctl and a linux/amd64 helper when the host is not Linux/amd64 (docs/plans/go-dev-launcher/).internal/agentctl/AGENTS.md— server routes, adapters, ACP;cmd/mock-agent/AGENTS.md— E2E scenario patterns and rebuild requirementsinternal/agentctl/server/api/AGENTS.md— reverse-proxy body rewriting (Accept-Encoding), iframe-blocking header strippinginternal/integrations/AGENTS.md— playbook for adding a new third-party integration (Jira/Linear pattern)docs/i18n.md("Backend") —internal/i18ncovers only what Go renders straight to a browser: the SPA-unavailable error pages and the shared-task artifacts (share.html, gist README and description). Both are complete. Everything else stays English by design; for new user-facing output prefer a stable error code the frontend translates. Usei18n.Tffor anything carrying a value — neverfmt.Sprintfa translated string, and never build a plural in Go. A locale for output that outlives the request is resolved once at write time and threaded as an argument, not a context value (ADR2026-08-01-share-artifact-locale.md).
