Imported from zohartito/sotto (
AGENTS.md). Install upstream withnpx skills add zohartito/sotto. Copyright stays with the author.
sotto — local, offline push-to-talk dictation: hold a key, speak, release; text lands at the cursor. macOS: PyObjC + mlx-whisper on Metal. Windows source alpha: sotto_win.py — pynput trigger hook (right-ctrl default), WASAPI capture, pinned faster-whisper (CUDA float16, CPU int8 fallback), paste-or-type insertion, pystray tray + Tk dialogs (--tray).
Run & test
All commands run from the repo root using the local venv (macOS only, Apple Silicon).
- Source alpha: native Python 3.12, macOS 14+ (dependency floor; only 27.2 tested).
python3.12 -m venv venv-alpha && venv-alpha/bin/python -m pip install -r requirements-alpha.txt; keep the sealed runtime dependency policy separate. - Isolated alpha verification:
venv-alpha/bin/python scripts/test_alpha.py(temporary data/cache; optional public calibration inputs require explicitSOTTO_CALIBRATION_FIXTURES). - Source launch: export
SOTTO_DATA_DIRandSOTTO_HF_HOMEto dedicated alpha folders; see README. Never use personal state for tests. - Permissions check:
venv/bin/python sotto.py doctor(Accessibility + mic) - Run:
venv/bin/python sotto.py run(trigger/hotkey come from Settings, default right Option;--trigger/--hotkeyoverride one run).sotto.py setup [--profile parakeet]pre-downloads pinned models + Silero VAD;sotto.py dictionarylists rules. - App:
scripts/install-mac.sh [--update|--uninstall]buildsSotto.appin/Applicationswhen writable, else~/Applications, removing our other copy (compiled launcher fromscripts/install_app.py, ad-hoc signed,SOTTO_LAUNCHER=app); login itemorg.sotto.alphavialogin_item.py(never the sealed label). - Full test suite:
venv/bin/python -m unittest discover -s tests(unittest; slow — loads heavy modules) - One module:
venv/bin/python -m unittest tests.test_speech_config(do NOT runvenv/bin/python tests/<file>.pydirectly — cwd isn't onsys.path, imports fail) - Windows PC suite:
venv\Scripts\python scripts\test_alpha.py(temp state + offline flags; plainunittest discover -s testswould use the real%APPDATA%\sotto— do NOT add-t; a global site-packagestestspackage shadows the repo'stests/dir). macOS-only tests self-skip on win32 (each skip carries a reason string). - Windows PC run:
venv\Scripts\python sotto_win.py(console;--trayfor the tray app;--device cpu|cudaorSOTTO_DEVICE;doctor,history,history-delete --id,history-clear --yes). Testers:scripts\install-windows.ps1(venv, pinned CPU/CUDA set, Start Menu shortcut viawin_launch.py+ pythonw;-WhatIf,-Uninstall) anddocs/windows-alpha.md(Python 3.13 x64;requirements-alpha-windows.txtCPU set + lock,requirements-alpha-windows-cuda.txtadds cuBLAS + its lock).scripts\sotto-win.batis the dev shortcut. - Gesture acceptance (8 scenarios):
venv/bin/python tests/gesture_check.py(wall-clock timed — run locally, flaky on CI runners) - Syntax-only check:
venv/bin/python -m compileall -q *.py tests
Structure
sotto.py— main entry + CLI (run/doctor/learning-*): GestureEngine, RecordingSession, transcription worker, cursor injection, LearningCoordinator/HistoryStore/LearningStore.speech_backends.py,speech_config.py,vad.py,ui.py— mlx Whisper/Parakeet backends & profiles (pins inMODEL_REVISIONS, 100WHISPER_LANGUAGES,automatic_languages), voice-activity detection (pinned Silero v6.2, verified on load), and the AppKit menu-bar + ember-orb overlay.- Engines in
sotto.py:LocalWhisper(one encoder pass for Automatic ≤30 s, experimental Fast = trimmed encoder,choose_languagenever translates clearly-other speech, unsure decodes and exhausted token budgets fall back to the full pipeline) andLocalParakeet(Parakeet TDT v3 on MLX, 25 European languages, array input, 120 s chunks); both resolve pinned snapshots viaresolve_pinned_snapshot(cache-first, local path only). - Shared pure modules (macOS + Windows):
dictionary.py(heard => write rules, correction suggestions),settings.py(settings.json),progress.py(History trends, no text). macOS UI:settings_window.py(PyObjC,@objc.python_methodhelpers). nemotron_backend.py,streaming_audio.py— optional, manually selected English streaming ASR using pinned NeMo-Speech.cpp Metal libraries; provision withvenv/bin/python scripts/setup_nemotron.py(Apple Silicon only). Pins live inconfig/nemotron-runtime.json; runtime only loads verified local files.sotto_win.py+win_asr.py/win_capture.py/win_hotkey.py/win_inject.py/win_ui.py(tray, Settings, Correct) /win_startup.py(HKCU Run login entry) /win_launch.py(pythonw entry) — Windows app (no overlay, no adaptive lane, no Nemotron). CT2 model ids live inwin_asr.PROFILES(+FAST_REPOfor Settings speed "fast"), exact commits inwin_asr.REVISIONS, per-repo file sets inwin_asr.REPO_FILES;speech_config.MODEL_PROFILESstill names the MLX repos for the Mac. Shareddictionary.py/settings.pyare Mac-owned APIs.- Adaptive/"silver" learning lane (headless
adaptive_worker.py):adaptive_learning.py,adaptive_runtime.py,silver_store.py,learning.py,history.py,teacher_backends.py,teacher_provision.py,calibration_worker.py,deployment_controller.py,sealed_release*.py. teachers/— Qwen3-ASR + Granite teacher adapters with isolated requirements;launchd/— plist templates;docs/— learning design notes;tests/— unittest suite.
Conventions & danger zones
- macOS runtime: Quartz/AppKit import at module scope (guarded in
sotto.pyso the pure logic imports on Windows; thewin_*modules raiseImportErroroff win32, and their tests skip with a reason). CI is deliberately tombstoned — see.github/workflows-disabled/README.md. Style: PEP 8, f-strings,pathlib. sotto_paths.pyis the only place that decides storage: macOS~/Library/Application Support/sotto(+Sotto/huggingface), Windows%APPDATA%\sotto(+huggingface);SOTTO_DATA_DIR/SOTTO_HF_HOMEoverride (blank = unset). Never add per-module platform path logic.- Windows ASR:
win_asr.resolve_model_dirruns once at startup — a complete cached snapshot of the pinned commit loads with no network call, a missing one downloads only when no offline flag is set, andWhisperModelonly ever gets that local directory. The first (warmup) transcription is the device probe: CTranslate2 loads cuBLAS lazily, so a CUDA failure surfaces there and falls back to CPU int8 unless--device cudawas forced. The idle rewarm runs only on CUDA (a CPU decode would queue ahead of the dictation).win_capturebinds each opened stream to its capture generation so a slow open can never leave the mic running; insertions run on one FIFO thread and wait while any modifier is held; nothing file- or icon-related runs on the keyboard-hook thread. Automatic = detect then argmax within the Settings languages (speech_config.automatic_languages), decoded with that language fixed, reusing the detection's encoder pass for clips up to 30 s (LanguageRestricted.last_language->preprocessing["detected_language"]). Dictionary viasotto.apply_personal_dictionaryonly on delivered text (live and retry);latency.release_to_text_secondsand the sharedprogressmodule match the Mac. Paste = delayed rendering from a clipboard-owner window (compare-and-restore on that thread after the foreground app reads it, retried if busy; owner-marked private content is typed instead; a tray-finished dictation is copied, not inserted); own injected keys carrywin_inject.INJECTED_TAGso the hook ignores them. One dictation app per Windows session (named mutex; tests setSOTTO_INSTANCE_KEY). Real-clipboard tests run in a private window station. - History Delete/Clear on Windows go through the same
sotto.nonadaptive_delete/nonadaptive_clearas the Mac menu.adaptive_learninglocks viastorage_lock.lock_exclusive(fcntl on macOS, msvcrt on Windows);ComparatorSpool.scrubwithout dir_fd authority (Windows) is complete only when the spool is absent. - Fully local/offline by design — no cloud, no telemetry, no speech upload. Inference must never resolve HuggingFace Hub ids. Startup resolves the exact commit pinned in
speech_config.MODEL_REVISIONScache-first (no network when that snapshot is cached), downloads it only on a first non-offline run, and passes inference a local path; offline flags forbid downloads. - The production LaunchAgent keeps only the selected ASR model resident. With
--idle-release 0, key-down starts the microphone engine/input tap and key-up stops/removes them; idle Sotto must not hold the microphone or retain idle audio. Whisper remains the default with sealedconfig/sotto-glossary.txt; optional Nemotron is a manual English-only engine, never an adaptive candidate or promotion. - The Speech engine submenu persists the choice and restarts only its own foreground process or supervising sealed LaunchAgent while idle. Nemotron (
--profile nemotron-en) streams canonical PCM on the single transcription worker, retains exact inference bytes for History/retry, and pastes finalized text only. It disables non-English/automatic language and Whisper glossary prompts; returning to Whisper restores the saved language choice. Never feed VAD-gated chunks or run MLX/native inference on the audio tap. Native load failure restores Whisper; native streaming failure retries the full canonical capture once. - Automatic language defaults to English only; users add their languages in Settings and each capture is detected among them. No language besides English is a default or special-cased. The menu-bar Language submenu persistently selects Automatic or one chosen language by changing only the decode hint; it must not load a second model or change the push-to-talk microphone lifecycle.
- After four idle minutes, the next baseline press schedules one single-flight model rewarm on the serialized transcription worker while capture is active; never move it to a concurrent MLX thread.
- Do not install the silver worker before public Calibration V2 and strict sealed readiness pass. Afterward it may collect shadow-only personal evidence; candidate routing remains impossible until personal evaluation authorizes at least 500 captures over at least 7 days.
- The learning system is gate-heavy on purpose (receipts, digests, revocation tombstones, one-shot promotion). Preserve those invariants; don't loosen them for convenience.
- Deploy ONLY via
scripts/rollout.sh(stage → activate → preflight under the new release → back up/install/reload the sealed LaunchAgent → verify, with automatic LKG + plist rollback and re-preflight on boot failure).kickstartalone does not reload changed ProgramArguments. Receipts are sealed toruntime_source_digest(), so any source change invalidates them; deploying by hand and skipping the preflight crash-loops the app at warmup ("adaptive baseline preflight receipt is invalid" — bit 4 rollouts before the script). - Sealed runtime markers must survive harmless cross-boot APFS
st_devrenumbering. Persist paths, inode/size/timestamps, exact tree membership, and SHA-256 bytes; usest_devonly for within-process race/alias checks, never as durable identity. - The ember overlay follows the display containing the pointer on every visible state and verifies WindowServer visibility after raising; preserve its one-shot panel rebuild and local diagnostic logging.
- The VAD is advisory only (trim + no-speech verdict AFTER transcription via
reads_as_no_speech). Never reintroduce an input-side gate that blocks captures from reaching the ASR: Silero scores real whispered dictation at 0% speech, and energy thresholds break on quiet whispering (2026-08-13/14 incidents). Exact-zero PCM is a dead mic route, not a quiet room (asr_skip_reason); it must not reach Whisper (2026-09-20: 41s of 24 kHz zeros →"ощ"x223). - Bit-exact zero PCM means THIS PROCESS's capture path is wedged, not that the room is quiet. 2026-09-20: AirPods reconnected under a new CoreAudio device id and every capture came back all-zero — including ones pinned to the built-in mic — while ffmpeg read those same AirPods at peak 18609 and a fresh process running
CaptureServicecaptured at peak 0.059. Stop/start does NOT clear it: theAVAudioEngineobject is reused for the process lifetime and carries the wedged HAL input unit across rebuilds._repair_dead_routedrops_engine_objfirst and relaunches via launchd second (guarded bySILENCE_RESTART_AFTER_Sso it cannot loop); the counter rearms only on real audio in_tap, never on engine start. Bluetooth warm-up emits ~410 ms of leading zeros, so the key-up rung needs >=1s of zeros before it fires. looks_hallucinatedis whole-transcript, so a whisper repetition loop that starts partway through used to cost the user the real dictation in front of it (2026-08-15: 403 good chars lost to"difference"x223).salvage_repetition_loopcuts at the loop and pastes the clean prefix, but only when that prefix passes the same guard on its own — keep it conservative; wall-to-wall garbage must still be quarantined, never pasted.- Sotto.app permissions (2026-09-30):
AXIsProcessTrusted()asks tccd once per process and then repeats that answer, so the app-mode dialog polls a fresh child (accessibility_granted_now) and relaunches after the grant. Sotto.app neverexecvs in place: same-pid exec leaves the new menu bar icon blank and unclickable (2026-10-01), so restarts go throughrelaunch_app_after_exit(detached helper waits for the launcher, then kickstarts the login item oropen -ns the app). The bundle is ad-hoc signed, so TCC pins the grant to that exact build ("Failed to match existing code requirement" in tccd logs);buildkeeps an identical bundle untouched — don't churn the launcher. Alerts come from Python, soui.use_app_iconsets Sotto's icon. - macOS disables session event taps silently (no callback notification) — 2026-09-15: process healthy, tap
enabled=False, 3 hours of dead hotkey.resync_pollerchecksCGEventTapIsEnabledevery second and callsrevive_tapDIRECTLY on the poller thread — never viaAppHelper.callAfter(the 2026-09-15 version did that and never fired; 2026-09-16 tap dead again). The menu bar has "Restart sotto" (bounded, drained exec for a foreground run; off-main-threadlaunchctl kickstart -konly when that job owns this PID; a process that looks sealed/launchd-started but cannot be confirmed refuses — never exec outside the bootstrap) because a clean Quit is a SuccessfulExit that KeepAlive does not relaunch. Diagnose withQuartz.CGGetEventTapList(look for the app pid with mask0x1c00). - GitHub Actions must stay SHA-pinned —
tests/test_workflow_action_pinning.pyenforces this. venv/,venv-alpha/,__pycache__/,.panel/,.handoff/are gitignored — never commit them.
Agent rules (all harnesses)
-
Determinism boundary: anything that must be EXACT lives in code; the model only orchestrates and judges.
-
Keep this file current: when you change how this repo is run, tested, or structured, update AGENTS.md in the same commit. Keep it under ~150 lines.
-
sotto_paths.pycentralizes alpha data/cache overrides without moving legacy data.launchd_templates.pyrenders home/FFmpeg placeholders during rollout; never install raw templates. -
Alpha tests do not install workers or enable adaptive routing. Run them from a checkout without world-writable or symlinked ancestors (not /tmp);
scripts/test_alpha.pyrefuses otherwise because the comparator-spool guards reject such paths.
