Architecture
WisprFree is a native macOS app; this site runs the same pipeline in the browser. Here's what runs where, what changes, and why.
The pipeline
Audio in
Speech → text
LLM cleanup
Text out
Stage 2 is optional in both builds. With no cleanup provider configured, the raw transcript is what you get — the app stays useful fully offline, and the web version degrades the same way.
Request flow
- 1
Capture
MediaRecorder collects Opus (Chrome/Firefox) or AAC (Safari) chunks while an AnalyserNode feeds an RMS level meter at ~20 Hz. Clips under 0.4 s are dropped as accidental taps — the same floor DictationPipeline uses.
- 2
Speech to text
The clip is uploaded to a route handler, which validates it, attaches the credential server-side, and forwards it to a hosted Whisper. Only the resulting text comes back — the credential never crosses into the browser.
- 3
Cleanup
The raw transcript, the active mode, and the dictionary go to Vertex AI behind a second handler, with a system prompt built by the same logic as PromptBuilder.swift. Thinking is set to MINIMAL — cleanup is a rewrite, not a reasoning task, and every millisecond shows.
- 4
Render and store
Both stages are TanStack Query mutations, so retries are off and the latency readout is the real round trip. The result is word-diffed against the raw transcript, then persisted. If step 3 fails, the raw transcript is shown instead — never a lost dictation.
Both stages are thin server-side proxies: they validate the payload, attach the credential, and hand back only the text. Because each call bills a real account, they are metered rather than left open.
Native vs. web
| Layer | macOS app | On the web |
|---|---|---|
| Shell | SwiftUI menu-bar app, AppKit overlay, XcodeGen project | Next.js 16 App Router, React 19, TypeScript, Tailwind v4, Motion, Lenis |
| Capture | AVAudioEngine → 16 kHz mono Float32 buffer | MediaRecorder (Opus/AAC) + AnalyserNode for the level meter |
| Speech to text | Parakeet TDT v2/v3, Whisper Large v3, or Cohere on CoreML — on-device | Groq-hosted Whisper large-v3-turbo, proxied through a route handler |
| Cleanup | Vertex AI, Gemini API, or any OpenAI-compatible endpoint | Vertex AI (gemini-3.6-flash) via @google/genai, same prompts |
| Secrets | macOS Keychain | Server-side only — the browser never sees a credential |
| Storage | JSON in ~/Library/Application Support/WisprFree | Zustand with the persist middleware, backed by localStorage |
| Delivery | Signed .zip + Sparkle auto-updates via an EdDSA-signed appcast | Static pages + serverless functions |
Privacy
In the macOS app, transcription happens entirely on-device. Audio never leaves your Mac. Only the cleanup step touches a network, and only if you configure a provider — your API key lives in the Keychain, not a config file. History, stats, and the dictionary are plain JSON under Application Support.
On the web, a browser can't host a 600 MB CoreML model, so the clip you record is POSTed to a serverless function which forwards it to Groq and returns the text. Nothing is written to a database, no session is created, and no audio is retained past the request. Your transcripts, settings, and dictionary live in this browser's localStorage — clear them any time.
That gap is the honest cost of running in a browser, and it's why the native app exists.