Imported from dabit3/speak (
AGENTS.md). Install upstream withnpx skills add dabit3/speak. Copyright stays with the author.
Speak
Speak is a native macOS 14+ app for dictation. It uses SwiftUI, AppKit, AVAudioEngine, and OpenAI GPT-Live-Transcribe. The Swift package has no external dependencies.
Use sans-serif fonts throughout the interface. Never introduce serif typography. Use the system default or rounded design, with monospaced digits only where needed for stable timers. Keep the dashboard free of the removed introductory tagline and status dot. Keep the Preferences Quit button in the fixed footer, outside the scrolling content. Separate that footer with a top border and keep native vertical scroll indicators enabled. Use normal app termination so microphone, shortcuts, and clipboard cleanup still run.
Run these commands from the project root:
- Run
swift testfor the automated tests. - Run
bash scripts/build.shto createbuild/artifacts.noindex/Speak.app. - Run
open build/artifacts.noindex/Speak.appto launch the app. - Run
bash scripts/package-dmg.shto package the existing signed app as a DMG. - Run
swift run -c release SpeakBenchmark --helpfor the live recognition benchmark.
The packaging script does not rebuild or sign the app again. It creates build/Speak-<version>-<architecture>.dmg and a SHA-256 checksum file. It refuses to overwrite an existing installer. The DMG contains Speak and an Applications shortcut. The script verifies the image, mounts it read-only, validates the bundled signature, and compares the executable with the source. Temporary packaging files stay under the hidden .build/ directory. App bundles stay under build/artifacts.noindex/ so Spotlight does not index development copies. The DMG also disables Spotlight indexing.
Keep only the latest verified installer and its checksum when cleanup is approved. After packaging a new version, ask for confirmation to delete specific older Speak DMGs and checksums. Eject only the older Speak images approved for cleanup. Leave source files, the installed app, the current app bundle, credentials, and unrelated disk images untouched.
The build script uses scripts/signing-identity.sh to select the only valid installed Developer ID Application certificate. If none or several exist, it stops and requires CODE_SIGN_IDENTITY. Certificate-signed builds enable the hardened runtime and a secure timestamp. Never silently fall back to an ad-hoc signature because its identity changes on each rebuild. CODE_SIGN_IDENTITY=- is an explicit temporary-development override and prints a permission warning. Keep the signing identity consistent across app updates. Switching from an ad-hoc build to a certificate-signed build requires a one-time permission approval. Release distribution also requires notarization through Apple, which signing alone does not perform.
Notarization uses the speak-notary Keychain profile. Never request or expose its password. For each public release, sign a staging copy of the DMG with the Developer ID certificate and a secure timestamp. Submit it with xcrun notarytool submit <dmg> --keychain-profile speak-notary and save the submission ID. Wait with xcrun notarytool wait <id> --keychain-profile speak-notary. Read the submission log even after acceptance. Only publish after Apple reports Accepted, xcrun stapler staple <dmg> and xcrun stapler validate <dmg> succeed, and Gatekeeper accepts both the DMG and its contained app. Assess the DMG with spctl --assess --type open --context context:primary-signature --verbose=2 <dmg>. Assess the contained app with spctl --assess --type execute --verbose=2 <app>. Replace the release artifact atomically and regenerate its SHA-256 checksum after stapling. Use a bare filename in the checksum so it works after download. Do not rebuild or re-sign accepted artifacts. Each new build requires its own notarization.
Public releases are GitHub releases on dabit3/speak with the stapled DMG and its .sha256 file attached. Keep the README download as a plain Download for macOS text link directly after the introductory paragraph, not an image button. It links to releases/latest/download/Speak-<version>-<architecture>.dmg, so update that filename in README.md after publishing each release. scripts/download-button.swift renders Resources/DownloadForMac.png from the shared logo file. Resources/ is not copied into the app bundle.
Speak 1.2.3 was accepted under submission 6769adff-86b2-4ec2-a6c6-02101f02d970 with no reported issues. The signed, stapled release is build/Speak-1.2.3-arm64.dmg, published as GitHub release v1.2.3. This approval does not cover later builds automatically. When you pass a commit to gh release create --target, use the full SHA from git rev-parse HEAD.
Sources/SpeakCore/SpeakLogo.swift defines the shared microphone-and-text-cursor mark. The interface, menu bar, icon generator, and SVG export use that path. The build script compiles scripts/icon.swift with the shared logo file and generates the app icon, Resources/Speak.png, and Resources/SpeakMark.svg. Keep the waveform bars only for the live audio meter, not for branding.
Sources/SpeakCore contains shortcut logic, API messages, and the transcription connection. Sources/Speak contains the native interface, microphone capture, Keychain access, and text insertion. Tests use a mock connection and do not call OpenAI.
To render native previews, create build/previews and run .build/debug/Speak --render-previews "$PWD/build/previews" after a debug build. The preview process does not request microphone or Accessibility access. AppKit renders the controls because SwiftUI ImageRenderer cannot render all native controls.
Use gpt-live-transcribe, not the conversational gpt-live-1 model or the older Whisper models. Stream mono 24 kHz PCM16 audio over a WebSocket, a persistent network connection. Disable automatic turn detection and commit audio after the user releases the shortcut. Match final transcripts to the committed item before pasting. The session sends the static TranscriptionConfiguration.prompt. It describes the setting and does not restate the transcription task, as OpenAI recommends. Keep it free of user data, line breaks, <, and >. The xhigh delay appears as "Most context." The default stays low because higher delays make the final transcript arrive later.
DictationFormatter runs locally on every partial and final transcript before smart correction. It applies English rules only when the language is English or auto-detect. It converts spoken punctuation and line-break commands, removes "um," "uh," and stutters, and writes spoken email addresses and domains. SpokenNumbers writes numbers of 10 or more, times, dates, years, money, percentages, decimals, and digit strings as digits. Standalone zero to nine stay words unless a measure, label, or range makes them numeric. Guards keep literal uses such as "the trial period," "Oxford comma," and "one day." Cover formatter changes in DictationFormatterTests. "Copy original dictation" returns the unformatted transcript. If formatting leaves no text, show the no-speech error and do not paste.
Keep API keys in macOS Keychain. Never put keys in source files, logs, or command arguments. Do not save audio or transcript history to disk. Keep only the last original and corrected transcripts in memory for copy and paste recovery.
Smart correction defaults to enabled. SmartCorrection debounces formatted partial text by 300 ms while recording, caps speculative requests, and reuses results only for identical source text. prepareFinal() runs when the user releases the shortcut. After that, each new partial is corrected after a 20 ms debounce, under a separate cap. The last partial usually matches the final transcript and arrives about 240 ms before it. Starting early lets most corrections finish within the wait limit. It waits at most 350 ms for final correction before falling back to the formatted text. When CorrectionPolicy.revisesItself finds a spoken self-correction cue, it waits at most 1 second. OpenAITranscriptCorrector uses gpt-4.1-nano-2025-04-14, predicted output, and store: false through Chat Completions. Predicted Outputs rejects max_completion_tokens, max_tokens, n, logprobs, tools, and penalties. Never add them. A rejected request fails silently, and Speak pastes uncorrected text. Speak 1.2.2 and earlier sent max_completion_tokens, so their corrections never ran. All corrections share OpenAITranscriptCorrector.session. prepare() warms that connection when recording starts. Fast mode (service_tier: "priority") did not reduce correction latency, so Speak does not use it. Its prompt fixes misheard words, self-corrections, disfluencies, punctuation, number formatting, and spoken email addresses. Keep prompt examples free of real names because the model copies them into transcripts. Adding more examples did not improve the benchmark and added about 50 ms. Transcript text, vocabulary, language, and the active app name are sent to OpenAI. No surrounding text or clipboard context is collected. Use CorrectionPolicy to reject unsafe edits. It converts number words to digits before it compares numbers, so an edit can change number format but not value. Without a self-correction cue, the edit must keep the same numbers, negations, dates, literals, names, and vocabulary, and it can add or remove only a few words. Capitalized words inside a sentence count as names. The first word of a sentence counts as a name only when Apple's NaturalLanguage tagger labels it as one, so a misheard first word can be fixed. The correction can replace any name with a vocabulary term. It can also respell a name slightly if the tagger does not label it as a person's name. The name and replacement must both have at least 5 letters, with 1 changed letter allowed, or 2 when either word has 8 or more letters. It cannot replace a name with a different name. A literal that matches a vocabulary term is compared without case. "a.m." and "p.m." are not literals. With a cue, the edit can remove those values and any number of words, but it cannot introduce new ones. Its checks are conservative heuristics, not a guarantee that an edit preserves meaning. Mock tests do not measure live model quality or response times.
SpeakBenchmark is a development tool that is not copied into the app bundle. It runs the real transcription, formatting, smart correction, and policy code against live OpenAI APIs. It reads the key from the OPENAI_API_KEY environment variable and never prints it. Ask the user before running it, because it adds API charges of about $0.06 per variant. It generates audio with macOS say voices, optionally mixed with noise, and streams it in real time. Audio and JSON results stay under .build/benchmark/. --replay <results.json> re-runs only correction on saved transcripts. Text-to-speech audio is not a real microphone, so use it to compare settings, not to claim live accuracy. In the September 2026 run, xhigh had about 24% fewer word errors than low but delivered the final transcript about 120 ms later at the median. Vocabulary hints cut word errors from about 5.5% to 3.8%.
showLiveTranscript defaults to true and is independent of showPill and smartCorrectionEnabled. Hidden live text must not stop audio capture, correction, or automatic paste. The compact pill must retain its recording controls and error messages. pillPosition defaults to bottom, which shows a horizontal pill with text above it. Left and right show a vertical pill halfway up that screen edge. Live text, errors, and success messages appear in a bubble beside the vertical pill, on the side that faces into the screen. Keep the vertical pill centered on the edge in every state. Keep its recording and finishing heights equal so it does not jump. Use Spinner instead of ProgressView in the dark pill so the progress indicator stays visible. The preview renderer writes a PNG for each side state.
The app needs microphone and Accessibility permissions for live dictation. Users enter their own API key in Preferences. A ChatGPT subscription does not include OpenAI API usage.
Paste automatically when dictation finishes. Capture the destination app when the user releases the shortcut. Do not require Accessibility element identity or metadata to send paste. Some text inputs do not expose that metadata. If the user changes apps while awaiting the final transcript, copy the text instead of pasting into another app. Block known password fields. Preserve the clipboard after automatic paste unless another process changes it. Never press Return or submit a form automatically.
Before claiming live accuracy or latency, test with a real microphone and an authorized API key. Automated tests do not establish live accuracy, actual latency, or compatibility with every target app. After rebuilding an app with a local signature, macOS can require new permission approval.
If Accessibility stays enabled in System Settings but Speak denies access, inspect macOS logs for a code requirement mismatch. A local rebuild can cause this mismatch. Get user approval before resetting permissions. Quit Speak, run tccutil reset Accessibility local.speak.dictation, and reopen the unchanged app. Ask the user to click Allow in Preferences and enable Speak again in System Settings. Do not rebuild between the reset and user approval. Do not reset other apps or other permission services.
Speak uses local.speak.dictation for its bundle identifier and Keychain service. A differently named predecessor does not share Speak preferences, permissions, or its Keychain entry. Ask the user to enter their API key and grant microphone and Accessibility permissions in Speak. Do not delete or read credentials from an older app during migration.
