Imported from elizaOS/eliza (
plugins/plugin-native-talkmode/AGENTS.md). Install upstream withnpx skills add elizaOS/eliza --skill plugin-native-talkmode. Copyright stays with the author.
@elizaos/capacitor-talkmode
Capacitor plugin for voice conversations: STT → chat orchestration → TTS, across browser, iOS, Android, and Electrobun (desktop).
Purpose / role
Provides a unified TalkMode Capacitor plugin that Eliza agents can call to run full voice conversation sessions. On native (iOS/Android/Electrobun), it uses platform STT and ElevenLabs streaming TTS with PCM/MP3 playback; on web it falls back to the Web Speech API for both STT and TTS. This is a Capacitor plugin, not an elizaOS Plugin object — it is imported directly into UI/app code via @elizaos/capacitor-talkmode, not registered through the elizaOS plugin registry.
Plugin surface
This is a Capacitor plugin exposing a single TalkMode object. It does not register elizaOS actions, providers, services, or evaluators. The surface is:
| Method / Event | Description |
|---|---|
TalkMode.start(options?) |
Start a voice session; accepts TalkModeConfig |
TalkMode.stop() |
Stop the voice session and release resources |
TalkMode.isEnabled() |
Query whether a session is active |
TalkMode.getState() |
Return current TalkModeState and statusText |
TalkMode.updateConfig(options) |
Patch config mid-session |
TalkMode.speak(options) |
Speak a string via TTS; returns SpeakResult |
TalkMode.stopSpeaking() |
Interrupt current TTS playback |
TalkMode.isSpeaking() |
Query TTS speaking status |
TalkMode.checkPermissions() |
Read microphone + speech-recognition permission status |
TalkMode.requestPermissions() |
Prompt for microphone + speech-recognition permissions |
Event: stateChange |
TalkModeStateEvent — state machine transitions |
Event: transcript |
TalkModeTranscriptEvent — interim and final STT results |
Event: speaking |
TTSSpeakingEvent — TTS utterance started |
Event: speakComplete |
TTSCompleteEvent — TTS utterance finished or interrupted |
Event: playbackStart |
TalkModePlaybackStartEvent — native PCM/MP3 playback started |
Event: error |
TalkModeErrorEvent — recoverable or fatal error |
Session modes (TalkModeSessionMode): idle, compose, push-to-talk, hands-free, passive.
State machine (TalkModeState): idle → listening → processing → speaking → error.
Layout
plugins/plugin-native-talkmode/
src/
index.ts Capacitor registerPlugin call; exports TalkMode singleton + all types
definitions.ts All TypeScript interfaces and types (TalkModePlugin, TTSConfig, etc.)
web.ts Web fallback: Web Speech API STT + SpeechSynthesis TTS
ios/
Sources/TalkModePlugin/
TalkModePlugin.swift Native iOS: AVSpeechSynthesizer + SFSpeechRecognizer + ElevenLabs PCM/MP3
android/
src/main/java/ai/eliza/plugins/talkmode/TalkModePlugin.kt Android native implementation (Kotlin)
ElizaosCapacitorTalkmode.podspec CocoaPods spec (requires AVFoundation + Speech frameworks)
rollup.config.mjs Builds IIFE (dist/plugin.js) and CJS (dist/plugin.cjs.js) from ESM
tsconfig.json
package.json
Commands
Scripts are defined in package.json; run them from the repo root with bun run --cwd:
bun run --cwd plugins/plugin-native-talkmode clean # remove build output
bun run --cwd plugins/plugin-native-talkmode build # build package artifacts
bun run --cwd plugins/plugin-native-talkmode build:docs # generate docs and build artifacts
bun run --cwd plugins/plugin-native-talkmode typecheck # TypeScript typecheck
bun run --cwd plugins/plugin-native-talkmode lint # mutating Biome check
bun run --cwd plugins/plugin-native-talkmode lint:check # read-only Biome check
bun run --cwd plugins/plugin-native-talkmode format # write formatting
bun run --cwd plugins/plugin-native-talkmode format:check # read-only formatting check
bun run --cwd plugins/plugin-native-talkmode test # run package tests
bun run --cwd plugins/plugin-native-talkmode prepublishOnly # publish-time build hook
bun run --cwd plugins/plugin-native-talkmode build:unlocked # bun run clean && tsc && bunx rollup -c rollup.config.mjs
bun run --cwd plugins/plugin-native-talkmode docgen # docgen --api TalkModePlugin --output-readme README.md --output-json dist/docs.json
Config / env vars
Config is passed at runtime via TalkMode.start({ config }) or TalkMode.updateConfig({ config }). No process-level env vars are read by this package. The key config fields:
| Field | Type | Notes |
|---|---|---|
tts.apiKey |
string |
ElevenLabs API key — required for ElevenLabs TTS on native |
tts.voiceId |
string |
ElevenLabs voice ID |
tts.modelId |
string |
ElevenLabs model (default on iOS: eleven_flash_v2_5) |
tts.outputFormat |
string |
e.g. "pcm_24000", "mp3_44100" |
tts.interruptOnSpeech |
boolean |
Stop TTS when mic detects speech |
tts.voiceAliases |
Record<string,string> |
Alias → voiceId mapping |
stt.engine |
"native" | "web" |
STT backend preference |
stt.modelSize |
"tiny" | "base" | "small" | "medium" | "large" |
Legacy compatibility field; ignored by current recognizers |
stt.language |
string |
BCP-47 language code (e.g. "en") |
stt.sampleRate |
number |
Audio sample rate in Hz (default 16000) |
silenceWindowMs |
number |
Silence gap before finalising transcript (ms) |
mode |
TalkModeSessionMode |
Initial session mode |
sessionKey |
string |
Chat session key passed to the orchestration layer |
The speak() call also accepts a TTSDirective for per-utterance overrides (voice, speed, stability, language, seed, etc.).
How to extend
Add a new method to the plugin surface:
- Declare the method signature in
src/definitions.tsonTalkModePlugin. - Implement it in
src/web.ts(web fallback). - Implement it in
ios/Sources/TalkModePlugin/TalkModePlugin.swift(register inpluginMethods). - Implement it in the Android Kotlin source at
android/src/main/java/ai/eliza/plugins/talkmode/TalkModePlugin.kt. - Run
bun run --cwd plugins/plugin-native-talkmode buildto verify TS compiles.
Add a new event:
- Define the event payload interface in
src/definitions.ts. - Add an
addListeneroverload toTalkModePlugininsrc/definitions.ts. - Call
this.notifyListeners("eventName", payload)in the web and native implementations.
Conventions / gotchas
- Not an elizaOS Plugin object. There is no
actions,providers,services, orevaluatorsarray. It is a Capacitor plugin registered withregisterPlugin("TalkMode", { web: loadWeb }). ImportTalkModefrom@elizaos/capacitor-talkmodein UI/app code. - ElevenLabs on web is blocked by CORS. The web implementation always falls back to
SpeechSynthesis;usedSystemTtswill always betruein the browser. ElevenLabs streaming TTS only works in native (iOS/Android) and Electrobun contexts. - iOS native frameworks required. The CocoaPods spec declares
AVFoundationandSpeechframeworks. iOS 13.0+ minimum deployment target. - Web STT auto-restarts.
recognition.onendrestarts the recogniser whenever the session is still enabled (this.enabled), preventing silent dropout when the browser ends a recognition run. Recognition is never paused during TTS, so the restart must not be gated onstate === "listening"; a spontaneousonendthat lands while the state is"speaking"(mid agent reply) would otherwise be swallowed, permanently deafening the single-shot recogniser (#22369). speak()on web always forceslangtoen-USunlessdirective.languageis set — this prevents browser-locale drift (e.g. numbers read in Chinese on Chinese-locale systems).- Silence detection is stateful. On iOS,
silenceWindow(default 0.7 s) drives aTasktimer that finalises in-flight transcripts. Adjust viasilenceWindowMsin config. - Peer dep:
@capacitor/core ^8.3.1is required at the app level. - See the root CLAUDE.md for repo-wide architecture rules, logger conventions, and git workflow.
Verification
Follow the repository-wide verification and evidence standard in the root CLAUDE.md. Run the package's relevant build, typecheck, lint, and test commands, then exercise the real integration boundary changed by the work. Inspect the produced domain artifacts and failure behavior; do not substitute mocked success for the system under test.