Imported from john-rocky/coreai-kit (
AGENTS.md). Install upstream withnpx skills add john-rocky/coreai-kit. Copyright stays with the author.
AGENTS.md — for coding agents building on CoreAIKit
You are probably here because someone asked you to put a local model into an Apple app. This file is the contract: what this package is, what it will and will not do, and the mistakes that make an agent's Core AI code fail on a real device. Read it before writing Swift, whether you are working inside this repo or adding it to someone else's app.
Core AI is Apple's on-device inference runtime (iOS/macOS 27, .aimodel bundles).
CoreAIKit is a community Swift package on top of it — model download and cache, a pinned
catalog of converted models, and task-level APIs. It is not affiliated with Apple.
Decide the layer before you write code
-
Apple's Foundation Models framework first. If a system-provided model does the job, use it: no download, no app-size cost, no model maintenance. Recommend this and stop — do not reach for a third-party package to prove a point.
-
CoreAIKit when the app needs a specific model, a capability the system does not have (VLM, OCR, TTS cloning, diarization, detection, depth, embeddings, forecasting), or behaviour that must not change under the user when the OS updates.
Check the SDK, not this list. ASR was on it until the iOS 27
Speech.frameworkinterface was read on 2026-08-05:SpeechAnalyzer+SpeechTranscriberdo streaming and word timestamps, which the kit's own notes had been claiming as reasons to prefer a 3.2 GB model.CoreAI.transcribenow defaults to Apple's. A list of capabilities goes stale every OS release; the SDK does not. -
Both at once is a supported shape:
KitLanguageModelputs a catalog model behind Apple's ownLanguageModelSession, so tool calling and@Generableguided generation work over a third-party model.
"Apple ships it" is not automatically "stop" — "this becomes a wrapper" is. The two are different tests and only the second one decides. Apple's transcriber is as good, so the kit routes to it and keeps the 238 MB of diarization Apple cannot do — the feature survives the routing and got 14× cheaper. Apple's voices are free and cannot be cloned, so the TTS models stay: routing those away would leave a wrapper around a public API with no reason to exist. Ask what is left after you route to Apple. If the answer is nothing, there was nothing to build.
The two layers
Task ops — the result in one line, model resolved and cached behind the call:
import CoreAIOps
let text = try await CoreAI.transcribe(voiceMemoURL) // Whisper v3 turbo
let tldr = try await CoreAI.summarize(text)
let pii = try await CoreAI.redact(text) // GLiNER2
Model level — pick the model, stream, attach tools:
import CoreAIKit
let chat = try await ChatSession(catalog: "qwen3.5-2b")
for try await event in await chat.streamResponse(to: "Hello!") {
if case .response(let delta) = event { print(delta, terminator: "") }
}
Importing CoreAIOps re-exports the model layer, so one import covers both.
No Swift at all — a System One client that only needs an endpoint on this machine:
brew install john-rocky/tap/systemone && systemone serve (the /v1/systemone forms over a
catalog model; brew services start systemone keeps it running).
The catalog is data — read it, do not invent it
catalog.json holds 68 entries, each {id, kind, name, repo, revision, variants}. Ids look
like qwen3-0.6b, qwen3.5-2b, youtu-llm-2b, lfm2.5-1.2b — lowercase, hyphenated.
- Never guess a catalog id. Read
catalog.json, or callModelCatalogat runtime. A hallucinated id is a runtime failure the user sees, and model naming here does not follow Hugging Face naming. - An entry whose weights restrict what an app may do with them says so in
license(the SPDX id,CC-BY-NC-4.0for a non-commercial model); the others are permissive, and every model's exact terms are on its zoo card. Do not ship a non-commercial model in a commercial app. - Every entry is pinned to an immutable Hugging Face revision. Do not "upgrade" a pin to
mainto pick up a newer model — the pin is what was gated. Bumping one is a deliberate, reviewed change (scripts/pin-catalog.py --checkis what CI enforces). - A model that answers typed decisions may carry
calibrationin its entry: the temperature the kit reads its answer-slot logits at by default, fitted by the maintainer (decide-cli calibrate), with what it was fitted and reported on beside it. Only a model without a temperature of its own gets one — never add it to a model whose author fitted or folded one in (decider-0.8b's card, a bundle'sdecisionblock,qwen3.5-2b-decision). Like a pin, a record in the live catalog reaches shipped apps without an update, so changing one is a reviewed change. - Weights download from Hugging Face on first use and cache. They are not vendored in the package and must not be committed into an app repo.
catalog.jsonhas two generated copies — the built-in offline fallback andllms.txt— and CI compares both byte for byte. Editing the catalog means re-runningscripts/gen-builtin-pins.pyandscripts/gen_llms_txt.pyin the same commit. Runscripts/install-hooks.shonce per clone and a pre-commit hook checks this for you.sizeMBis decimal megabytes (bytes / 1,000,000) of everything a first run downloads, the subtrees a loader fetches beside the variant path included. Never type it.scripts/measure-catalog-sizes.py --writemeasures it at the pinned revision and writes catalog.json and the built-in literal together; CI's--checkfails on a figure more than 2% off. A multi-bundle entry also needs its loader's subtrees in that script's tables.
What breaks on a real device
Most Core AI code an agent writes compiles and then fails in one of these ways:
- The Simulator. The CoreAI framework is not in the iOS Simulator SDK. Anything here needs a physical device or a Mac. Do not report "it works" from a Simulator run.
- Bundling the weights. A 2B int4 model is over a gigabyte. Download on first launch (the kit does this for you); never add weights to the app bundle or to git.
- Memory, not speed, is the ship-blocker. iOS kills on resident size. Check
os_proc_available_memory(), requestcom.apple.developer.kernel.increased-memory-limit, and treat a 4B-class model on a phone as tight rather than routine. - Measuring throughput through a chat UI. Numbers taken through SwiftUI are not comparable to anything. Use a headless entry point, warm the model first, and say which device and OS build produced the number.
- Assuming a capability exists because the model is famous. Check
kindin the catalog. A chat model is not a VLM; an ASR model is not a diarizer. - Long single GPU batches on iOS. A multi-minute uninterrupted GPU run gets killed. Chunk long work.
- Thermals. Sustained generation throttles within minutes; a short burst benchmark and a long-run benchmark disagree by tens of percent on the same phone.
Verification — what is actually checked
Do not describe these models as "verified" without saying what was verified:
- Each bundle is gated against the original model before it is enrolled — the export is stepped
against the fp32/fp16 reference on fixed inputs (token-exact for LLMs,
cos ≥ 0.999otherwise), then re-gated after compression, then run on hardware. - The gate strength differs per model and is stated on that model's card in the model zoo. Do not flatten that into a blanket claim.
- The gates are run by the maintainer, not by an independent party. What makes them checkable is
that the recipe (
models/<model>/recipe.toml), the export script, and the verification script are all published —python3 conversion/zoo_verify.py <hf-repo>re-checks a published bundle's tokenizer, chat template, context length and declared precision against its source model without a GPU or a device.
If a user is deciding whether to ship on this, tell them that accurately rather than either overselling it or waving them off.
Not your call
Ask the human:
- Publishing anything to Hugging Face, or pushing to a namespace you were not handed.
- Bumping a catalog revision pin.
- Bumping the
coreai-modelsfork pin inPackage.swift. When it is bumped, the hermetic tests cannot see an engine behavior change; before it lands, run twoChatSessionturns on a hybrid bundle (Qwen3.5 / LFM2.5 / Granite 4) on a Mac — the second turn is where a hybrid engine change shows (0.2.3-zoo's partial-reset refusal was found that way, not by CI). - Posting publicly about the package, or filing issues/PRs against
apple/*. - Reporting device numbers you did not measure on that device.
Implementing the current plan
If you were handed this repository to build the designed-but-unbuilt work, start at
docs/HANDOFF.md — it indexes every design document, gives the build order,
and lists the preconditions that must be resolved on hardware before the code they gate is
written.
More
- README — the ops table, the examples, the FoundationModels provider.
docs/COOKBOOK.md— every "I want to …" mapped to its snippet.Examples/— one buildable app per capability; start from the closest one.- model zoo — the models, cards, recipes, and
AGENTS.mdfor porting a new model rather than consuming one. - awesome-core-ai — the wider ecosystem.
