Speech processing
Where microphone audio becomes a transcript. This is narrower than later cleanup, formatting, assistant, or destination processing.
Open research · July 2026
“Local,” “cloud,” and “zero retention” answer different questions. This source-reviewed matrix separates speech processing, optional connected features, offline use, retention, accounts, and platform requirements for 17 current Mac dictation products.
Version 1.2.0 · CC BY 4.0 · cite the original sources
Read the labels precisely
The matrix does not collapse different architectures into a badge or winner. These definitions keep each claim inside its documented boundary.
Where microphone audio becomes a transcript. This is narrower than later cleanup, formatting, assistant, or destination processing.
Whether the product exposes an optional provider, hosted model, assistant, cleanup, or agent path that can change what leaves the Mac.
Whether the documented speech-to-text path can work without a live connection after required model, app, or license setup.
What the publisher says it or a named processor keeps. “Not retained as a recording” does not mean no processing or no resulting text history.
Whether an account, subscription, license check, provider key, or synced history participates in the documented workflow.
The matrix
Scroll horizontally to inspect every boundary →
| Product | Speech processing | Configurable cloud path | Offline availability | Retention boundary | Account boundary | Platform | What this does not prove |
|---|---|---|---|---|---|---|---|
| IraVoice1 | Speech recognition and supported formatting run on the Mac after model setup. | No cloud speech mode. Optional direct agent handoff sends the finished specification, not microphone audio. | Dictate works after the speech-model download; paid continuation works after one Gumroad key check. | IraVoice says its servers receive no audio, transcript, selected text, vocabulary, or repository context; the Mac app has no product telemetry or cloud transcript history. | No account or card for the full trial. Paid use relies on a Gumroad license key stored in Keychain. | macOS 26+ on Apple silicon. | IraVoice publishes this matrix. Verify its product claims against the linked privacy page and treat the dataset as documentation research, not an independent audit. |
| Apple Dictation23 | On-device or on Apple servers depending on language and the indicator shown in Keyboard Settings. | The processing path is configuration-dependent rather than a separate third-party provider selector. | Available without a connection only for configurations Apple marks as on-device. | Apple says server-processed Dictation is not stored unless Improve Siri & Dictation is enabled; request and transcript metadata have separate documented handling. | Included with supported macOS. Apple documents device-generated identifiers and separate Apple Account-related service boundaries. | Built into supported macOS versions. | Do not assume every language or Mac configuration uses the same processing path. Check the current Keyboard Settings disclosure on the actual Mac. |
| Google AI Edge Eloquent45 | Google’s reviewed Mac launch sources describe the current native macOS feature set as on-device, including dictation, Voice Edit, and local audio or video file transcription. | The current Mac product page says Google AI Edge Eloquent can optionally access Google Workspace data such as Gmail to generate a vocabulary list of unique words. | Google’s June 3, 2026 launch announcement says the current native macOS feature set is fully offline. | The reviewed Mac sources establish on-device processing and offline use, but do not publish a separate desktop retention or telemetry statement for every Mac feature. | The reviewed Mac sources document a direct Mac download and optional Workspace or Gmail vocabulary access, but do not clearly state whether a Google account is otherwise required for the core desktop workflow. | Native macOS app; the current FAQ says English only. | The current product-page FAQ still contains older desktop-evaluation wording that conflicts with the dated launch announcement and direct Mac DMG. Use the launch post and current Mac download as platform truth, and do not transfer iOS-only cloud-feature wording to the Mac app. |
| BetterDictation678 | Initial speech recognition uses local Whisper on a compatible Apple-silicon Mac. | Optional Pro sends the text transcript—but not audio—to OpenAI for cleanup and formatting. | Basic local transcription works offline. Pro cleanup requires a connection. | The privacy policy separates local audio processing from optional transcript transfer and also discloses performance and crash telemetry. | A paid license key is required; optional Pro is a separate recurring service. | Apple-silicon Mac; current help asks for macOS Sonoma 14.4 or later. | The local claim covers initial transcription. Enabling Pro changes the transcript path even though microphone audio remains local. |
| Spokenly9101112 | Local Whisper or Parakeet, user-key providers, or managed hosted models depending on the active mode. | Yes. Model, AI Instructions, provider, optional context, and output actions can change the data path. | Local models work offline. Local Only Mode blocks outbound access while permitting local and localhost models. | Retention depends on local, managed, or bring-your-own provider paths; Spokenly also documents history and optional audio replay. | Local and BYOK paths are available without a paid Pro plan; managed services and sync features can introduce account boundaries. | macOS, Windows, Linux, and iOS, with features varying by platform. | “Spokenly” is not one fixed architecture. Record the active mode, provider, context sources, history setting, and output action. |
| Raycast Dictation1314 | Raycast says recorded audio is sent to its speech-to-text partner for transcription. | Dictation is a connected Raycast AI feature; optional App Context and later AI actions can add separate data. | No documented offline speech-to-text path for Raycast Dictation. | Raycast says it does not retain recorded audio after the transcription request; Dictation history is stored locally on the Mac. | Raycast AI plan and account controls apply to the connected feature. | Raycast for macOS on supported Macs. | No post-request audio retention does not mean on-device transcription. Local history and any later AI action are separate boundaries. |
| Wispr Flow15 | Wispr says transcription always happens in the cloud. | The speech path is cloud-based. Privacy Mode and Private Cloud Sync settings change retention, not the location of transcription. | No documented offline transcription path. | Wispr says Privacy Mode with Private Cloud Sync disabled provides zero retention; other settings can retain data for history or sync. | A connected product account and plan participate in the service. | Current desktop and mobile platforms supported by Wispr Flow. | Zero retention is a storage policy for the cloud path, not a claim that microphone audio never leaves the device. |
| Superwhisper1617 | Superwhisper documents offline, on-device transcription for its local speech path. | Optional cloud models and browser or assistant tools are separate connected paths. | The documented local transcription path works offline after required model setup. | For offline transcription, Superwhisper says audio is not uploaded. Connected features follow their selected service boundaries. | App licensing applies; optional connected providers or tools can introduce separate credentials and terms. | Mac, Windows, and iPhone support is currently advertised. | The offline claim belongs to local transcription. Do not transfer it to an optional cloud model, browser tool, or assistant feature. |
| MacWhisper1819 | MacWhisper documents local transcription by default, including a local dictation path. | Optional cloud transcription, assistant, and third-party AI provider features can change the path. | Supported local models and dictation can work offline after model setup. | Local workflows keep transcription on the Mac; retention for optional providers depends on the chosen external service. | Free and paid app tiers exist; optional provider features can require separate keys or accounts. | macOS, with model performance depending on Mac hardware. | MacWhisper includes broader file, meeting, assistant, and provider workflows. The active model and feature determine the actual boundary. |
| VoiceInk202122 | Local models are the default; users can configure cloud transcription and enhancement providers. | Yes. The active mode can select cloud transcription, enhancement, context, auto-send, or a custom command. | Local model workflows can work offline; configured cloud providers require a connection. | Local defaults keep speech processing on the Mac. Optional provider retention follows the provider selected by the user. | GPLv3 source can be self-built; the compiled app uses a commercial license, and optional providers use user credentials. | macOS 14.4 or later. | The active model, provider, context source, mode, and output command determine the boundary; an open-source codebase does not make every configuration local. |
| Voice Type23 | Core dictation, custom vocabulary, formatting, and file transcription run on the Mac. | Optional rewrite providers require a connection and send selected text to the provider the user chooses. | Core dictation works offline after required app and model setup; optional rewrites do not. | The App Store listing describes on-device core processing. Apple may receive purchase and receipt-verification traffic. | Distributed through the Mac App Store, so Apple purchase and receipt boundaries apply. | macOS 13.4 or later. | The on-device claim covers core dictation. Enabling a rewrite provider changes the selected-text path, and App Store licensing still creates an Apple network boundary. |
| Whryte24 | The publisher documents local Parakeet transcription and local smart formatting. | No cloud speech or formatting provider is disclosed for the documented workflow. | The documented speech path works offline after the model download. | History is stored locally; the publisher says audio and transcripts are not transmitted. | No product account is required. Purchase and license delivery use Gumroad. | macOS 14 or later on Apple silicon. | This is a review of publisher documentation. A “100% offline” statement does not independently verify the current binary or every licensing request. |
| TypeVox25 | WhisperKit runs locally on Apple silicon; the publisher says audio is processed in memory. | No cloud speech or cleanup path is documented for the current core product. | The documented speech path works offline after model setup. | The publisher says audio is never saved or transmitted. | No product account is required; Pro uses an emailed license key. | macOS 15 or later on Apple silicon. | The product page does not establish an independent audit or guarantee that every future feature shares the same boundary. |
| Dicta2627 | Speech recognition runs locally using a model downloaded to the Mac. | The documented current product does not offer an AI rewrite or cloud speech path. | Dictation works offline after the roughly 3 GB model download. | The app does not upload audio or transcripts; the privacy policy separately discloses account, license, device, billing, and consented website-analytics data. | The website version uses an account and recurring license; Apple handles payment and account boundaries for the App Store version. | macOS 14 or later; Intel Macs are supported with slower documented performance. | The app-level no-analytics claim should not be transferred to the website, account, license, billing, or optional analytics boundaries. |
| MacParakeet2829 | Speech recognition runs locally using supported on-device engines. | Optional AI features can use local providers or send transcript text—not microphone audio—to a selected cloud provider. | Local speech recognition works offline after required model setup. | The publisher says it does not collect audio or transcript content; optional anonymous telemetry can be disabled, and meeting retention is controlled locally. | No product account is required for the documented app workflow. | macOS 14.2 or later on Apple silicon. | The local speech claim does not extend to a cloud AI provider the user selects or to externally downloaded source content. |
| Susurr30 | WhisperKit transcription runs locally on Apple silicon. | No cloud dependency is documented for the current transcription workflow. | Transcription works offline after required model setup. | The product page says transcription history is stored locally. | A one-time Polar license key can activate up to three Macs; no recurring core plan is advertised. | macOS 14 or later on Apple silicon. | The public product page does not document an independent audit or every licensing and network-metadata boundary. |
| Sotto31 | Users can select local WhisperKit or Parakeet models, or optional OpenAI and Groq cloud transcription with their own key. | Yes. Cloud mode sends recordings directly to the selected provider; the publisher says no Sotto server sits in the middle. | Local models work offline after download. Cloud providers require a connection. | Local-mode audio stays on the Mac. Cloud-mode retention and processing follow the selected provider, while saved history remains a separate local setting. | Purchase uses an emailed license key for up to three Macs; optional cloud providers use the user’s credentials. | macOS 13 or later, with Apple silicon recommended. | The active local or cloud model determines the speech boundary. Cloud-provider retention and local history must be evaluated separately. |
Source numbers in the product column jump to the reviewed first-party documentation below. Policies and features can change; the date above is part of the data.
What the matrix reveals
A product can transcribe locally and still offer cloud cleanup, assistant features, sync, or an agent handoff. Inspect the active mode rather than inheriting a local claim from one feature.
A cloud service may process speech in real time while promising not to retain audio. That is a different architecture from keeping microphone audio on the Mac.
After dictation inserts text, the receiving editor, browser, CRM, chat, or AI agent can store or process it under separate terms. This matrix stops at the dictation product boundary.
Method and limitations
The dataset is intentionally reviewable. Each row points back to the publisher’s current documentation and includes a limitation that should travel with the claim.
Primary sources
Reviewed July 26, 2026. Product versions, settings, policies, and availability can change after this date.
Local speech and formatting, required connections, telemetry, license, and optional agent-handoff boundaries.
Configuration indicator, supported workflow, settings, and on-device availability.
On-device and server processing, Improve Siri & Dictation, identifiers, and retention disclosures.
June 3, 2026 native macOS launch, fully offline on-device processing, Voice Edit, local file transcription, and Gemma 4 12B details.
Current Mac download, no-cap and no-subscription positioning, English-only FAQ note, optional Workspace or Gmail vocabulary access, and stale desktop FAQ wording caveat.
Local Whisper, languages, activation modes, platform requirements, and license.
Optional OpenAI transcript cleanup and connected pricing.
Local initial transcription, Pro transcript transfer, audio boundary, and telemetry.
Local and hosted models, platforms, activation, and agent features.
Providers, AI Instructions, context, triggers, and output actions.
Local, managed, BYOK, account, analytics, context, and provider boundaries.
Outbound blocking, localhost allowance, supported models, and failure behavior.
Speech-to-text workflow, history, App Context, and current controls.
Speech partner processing and post-request audio retention statement.
Cloud transcription, Privacy Mode, Private Cloud Sync, and zero-retention configuration.
On-device speech path, offline availability, upload boundary, and separate connected tools.
Current platforms and local or cloud feature positioning.
Local default, model and provider options, app tiers, and broader transcription features.
Local dictation path, offline use, and hardware guidance.
Local and cloud models, modes, context, shortcuts, and output behavior.
Dual licensing, compiled-app license, local defaults, and optional provider boundaries.
Provider, context, auto-send, output, and custom-command configuration.
On-device core features, optional rewrite-provider path, offline use, platform requirement, and App Store boundaries.
Local Parakeet processing, offline use, local history, account boundary, model download, and platform requirements.
Local WhisperKit processing, in-memory audio handling, offline use, license boundary, and platform requirements.
Local model, offline use, model download, no rewrite path, and Intel Mac support.
Local speech processing, audio and transcript boundaries, account and license metadata, billing, and optional website analytics.
Local speech recognition, optional AI providers, content collection, telemetry, account, and platform requirements.
Local speech engines, local and cloud LLM paths, meeting retention controls, telemetry, and architecture details.
Local WhisperKit transcription, offline use, local history, license boundary, and platform requirements.
Local and BYOK cloud model paths, direct-provider boundary, offline use, license, history, and platform requirements.
Questions
The reviewed sources document local speech paths for IraVoice, BetterDictation, Spokenly local models, Superwhisper local transcription, MacWhisper local models, VoiceInk local models, Voice Type core dictation, Whryte, TypeVox, Dicta, MacParakeet, Susurr, and Sotto local models. Apple Dictation is configuration-dependent. Always check the active mode and current product documentation.
No. Zero retention describes what a service says it keeps after processing. A product can transcribe speech in the cloud and promise zero retention, or transcribe locally and later send the text to a cloud cleanup or destination service.
Yes. Optional cleanup, assistant, provider, sync, context, or agent features can create a connected path after local speech recognition. This matrix lists those paths separately.
No. It is a dated review of first-party public documentation. It does not inspect network traffic, verify source code, test every mode, or certify compliance. Follow the linked sources and verify the active configuration before sensitive use.
The downloads make the same dated rows easier to inspect, cite, diff, and reuse in research or AI systems without scraping the visual table. They do not grant permission to remove source attribution or product caveats.
Test the actual configuration
If IraVoice’s local, Apple-silicon workflow fits your constraints, try it for 14 days with no card or account.