Architecture

WisprFree is a native macOS app; this site runs the same pipeline in the browser. Here's what runs where, what changes, and why.

The pipeline

01

Audio in

→
02

Speech → text

→
03

LLM cleanup

→
04

Text out

Stage 2 is optional in both builds. With no cleanup provider configured, the raw transcript is what you get — the app stays useful fully offline, and the web version degrades the same way.

Request flow

  1. 1

    Capture

    MediaRecorder collects Opus (Chrome/Firefox) or AAC (Safari) chunks while an AnalyserNode feeds an RMS level meter at ~20 Hz. Clips under 0.4 s are dropped as accidental taps — the same floor DictationPipeline uses.

  2. 2

    Speech to text

    The clip is uploaded to a route handler, which validates it, attaches the credential server-side, and forwards it to a hosted Whisper. Only the resulting text comes back — the credential never crosses into the browser.

  3. 3

    Cleanup

    The raw transcript, the active mode, and the dictionary go to Vertex AI behind a second handler, with a system prompt built by the same logic as PromptBuilder.swift. Thinking is set to MINIMAL — cleanup is a rewrite, not a reasoning task, and every millisecond shows.

  4. 4

    Render and store

    Both stages are TanStack Query mutations, so retries are off and the latency readout is the real round trip. The result is word-diffed against the raw transcript, then persisted. If step 3 fails, the raw transcript is shown instead — never a lost dictation.

Both stages are thin server-side proxies: they validate the payload, attach the credential, and hand back only the text. Because each call bills a real account, they are metered rather than left open.

Native vs. web

LayermacOS appOn the web
ShellSwiftUI menu-bar app, AppKit overlay, XcodeGen projectNext.js 16 App Router, React 19, TypeScript, Tailwind v4, Motion, Lenis
CaptureAVAudioEngine → 16 kHz mono Float32 bufferMediaRecorder (Opus/AAC) + AnalyserNode for the level meter
Speech to textParakeet TDT v2/v3, Whisper Large v3, or Cohere on CoreML — on-deviceGroq-hosted Whisper large-v3-turbo, proxied through a route handler
CleanupVertex AI, Gemini API, or any OpenAI-compatible endpointVertex AI (gemini-3.6-flash) via @google/genai, same prompts
SecretsmacOS KeychainServer-side only — the browser never sees a credential
StorageJSON in ~/Library/Application Support/WisprFreeZustand with the persist middleware, backed by localStorage
DeliverySigned .zip + Sparkle auto-updates via an EdDSA-signed appcastStatic pages + serverless functions

Privacy

In the macOS app, transcription happens entirely on-device. Audio never leaves your Mac. Only the cleanup step touches a network, and only if you configure a provider — your API key lives in the Keychain, not a config file. History, stats, and the dictionary are plain JSON under Application Support.

On the web, a browser can't host a 600 MB CoreML model, so the clip you record is POSTed to a serverless function which forwards it to Groq and returns the text. Nothing is written to a database, no session is created, and no audio is retained past the request. Your transcripts, settings, and dictionary live in this browser's localStorage — clear them any time.

That gap is the honest cost of running in a browser, and it's why the native app exists.