On-device speech-to-text benchmark: Parakeet vs Whisper vs Gemma, tested on a laptop
We ran every speech model Walkie ships — plus Gemma 4's audio mode — on one ordinary laptop, then checked the results against the best-known public benchmark of real human speech. Here is what we found, what we run, and why.
Short answer. For dictation on your own computer, NVIDIA's Parakeet is the best default: it transcribed a typical clip in 0.16 s on a base M4 laptop — 19 times faster than Whisper Large V3 Turbo — and no model on the Open ASR Leaderboard's real-world English that is licensed for commercial use is both faster and more accurate. Whisper Large V3 Turbo, Walkie's other main option, is more accurate on rare words and covers 99 languages, at about three seconds per clip. The best paid cloud engines, from Zoom, Microsoft and ElevenLabs, beat everything that runs locally — if you are willing to send them your audio.
What Walkie runs
Walkie has two modes. On-Device Mode runs a speech model on your computer and makes no network request. Cloud Mode sends the audio to a transcription provider and adds an AI cleanup step for punctuation, filler words and formatting. You switch between them whenever you like.
| Model | Maker | Download | Languages | Where in Walkie | Pick it for |
|---|---|---|---|---|---|
| Parakeet V3 | NVIDIA | 478 MB | 25 European | On-device default | The default. Fast and accurate in 25 languages |
| Whisper Small | OpenAI | 487 MB | 99 | On-device option | Other languages on an older computer |
| Whisper Medium | OpenAI | 492 MB | 99 | On-device option | More accuracy than Small, 99 languages |
| Whisper Large V3 Turbo | OpenAI | 1.6 GB | 99 | On-device option · Cloud Mode | The most accurate option, 99 languages; also Cloud Mode's engine |
| Moonshine V2 Medium | Useful Sensors | 192 MB | English | On-device option | US English on modest hardware |
| Moonshine V2 Small | Useful Sensors | 100 MB | English | On-device option | US English, smaller still |
| Moonshine V2 Tiny | Useful Sensors | 31 MB | English | On-device option | The smallest download in Walkie |
We retired five models this year. SenseVoice, Breeze ASR and Moonshine Base were almost never chosen; Parakeet V2 and Whisper Large are covered by V3 and Turbo, which the tests below explain. Anyone who already has one keeps it. A shorter list of models that each earn their place beats a long one you have to research.
How to read the numbers
Word error rate is the standard accuracy score: out of every 100 words, how many came out wrong — swapped, dropped or added. 5% means one mistake in twenty words. Lower is better. Both our test and the leaderboard strip punctuation and capitals before scoring, so a model is not penalised for writing “4:30” where another writes “four thirty”.
Speed is shown two ways. Our test reports the real time a laptop took per clip, which is what you feel. The leaderboard reports how many seconds of audio a model gets through per second on a data-centre GPU — useful for comparing models with each other, not for predicting your laptop.
Which on-device model is fastest? Every model, one laptop
We read 20 realistic dictation phrases in 4 English accents — 80 clips, 7 minutes of audio — through each model on a 14-inch MacBook Pro, Apple M4 (10-core CPU, 10-core GPU), 24 GB. Every model ran on the same engine builds the Walkie app uses, loaded once and warmed up once, the way the app keeps a model ready between dictations.
| Model | Errors | Time per clip | Faster than real time | Download | In Walkie |
|---|---|---|---|---|---|
| Whisper Large V3 Turbo | 2.2% | 2.9 s | 1.7× | 1.6 GB | On-device option · Cloud Mode |
| Whisper Large V3 | 2.4% | 4.0 s | 1.3× | 1.1 GB | — |
| Gemma 4 E4B (audio) | 2.4% | 1.3 s | 3.8× | 5.2 GB | — |
| Parakeet V2 | 3.1% | 0.16 s | 32.1× | 473 MB | — |
| Whisper Medium | 3.1% | 2.0 s | 2.4× | 492 MB | On-device option |
| Whisper Small | 3.5% | 0.75 s | 6.6× | 487 MB | On-device option |
| Gemma 4 E2B (audio) | 3.7% | 0.68 s | 7.2× | 3.6 GB | — |
| Parakeet V3 | 4.4% | 0.16 s | 32.6× | 478 MB | On-device default |
| Moonshine V2 Medium | 10.8% | 0.35 s | 14.1× | 192 MB | On-device option |
| Moonshine V2 Small | 12.6% | 0.24 s | 20.8× | 100 MB | On-device option |
| Moonshine V2 Tiny | 25.0% | 0.11 s | 43.9× | 31 MB | On-device option |
| Model | US | UK | India | Australia |
|---|---|---|---|---|
| Whisper Large V3 Turbo | 2.1% | 2.4% | 1.7% | 2.4% |
| Whisper Large V3 | 2.1% | 2.4% | 2.4% | 2.8% |
| Gemma 4 E4B (audio) | 2.4% | 2.1% | 2.8% | 2.4% |
| Parakeet V2 | 3.5% | 2.8% | 2.8% | 3.1% |
| Whisper Medium | 2.8% | 1.7% | 3.5% | 4.2% |
| Whisper Small | 3.1% | 2.4% | 4.9% | 3.5% |
| Gemma 4 E2B (audio) | 2.1% | 4.2% | 4.9% | 3.5% |
| Parakeet V3 | 2.1% | 4.5% | 6.3% | 4.5% |
| Moonshine V2 Medium | 1.7% | 16.1% | 12.2% | 12.9% |
| Moonshine V2 Small | 3.8% | 17.5% | 16.4% | 12.6% |
| Moonshine V2 Tiny | 9.1% | 23.1% | 25.2% | 42.7% |
What stood out
- Parakeet is in a speed class of its own. Both versions took about 0.16 s per clip — around 33 times faster than real time. Whisper Large V3 Turbo took 2.9 s and Whisper Large V3 took 4.0 s. Large was slower than Turbo without being more accurate here, which is why Walkie offers Turbo instead. When you dictate dozens of times a day, that gap is the difference between text that appears and text you wait for.
- Whisper is the most accurate on clean audio, at 2.2% for Turbo. Parakeet V2 made 3.1% errors, tying Whisper Medium while running 13 times faster. V2 is English-only, so Walkie ships V3: the same speed, 24 more languages, and about a point more error in English.
- Parakeet's misses were mostly jargon. “Kubernetes” came out as “Cuban Eats” and “amoxicillin” was misspelled; Whisper got both. This is what Walkie's custom dictionary is for: add the word once and it comes out right in either mode.
- Whisper Small is the middle ground — 3.5% errors in 0.75 s, in 99 languages.
- Moonshine is built for US English. Moonshine V2 Medium scored 1.7% on the US voice and 16.1% on the British one; Tiny reached 42.7% on the Australian voice.
Computer voices are cleaner than people. They have no background noise, no mumbling and no crosstalk, which flatters models with strong language knowledge — Whisper and Gemma most of all. Treat the speed column as the reliable result here and check accuracy against the leaderboard below, which uses real recordings. One more note: Parakeet writes dollar amounts without a space before the sign (“is$25,000”). Walkie puts the space back before pasting, and the scores above include that fix.
A real voice through a laptop microphone
Computer voices are a fair way to compare speed and an unfair way to judge accuracy. So we ran the same 20 phrases again, read aloud by one person into a MacBook Pro built-in microphone, trimmed the way Walkie trims a dictation.
| Model | Errors, real voice | Errors, computer voices | Time per clip | In Walkie |
|---|---|---|---|---|
| Whisper Medium | 12.9% | 3.1% | 2.0 s | On-device option |
| Whisper Large V3 Turbo | 12.9% | 2.2% | 2.9 s | On-device option · Cloud Mode |
| Parakeet V3 | 13.6% | 4.4% | 0.13 s | On-device default |
| Apple SpeechAnalyzer (macOS) | 15.0% | — | 0.10 s | — |
| Gemma 4 E4B (audio) | 16.1% | 2.4% | 1.4 s | — |
| Gemma 4 E2B (audio) | 16.8% | 3.7% | 0.68 s | — |
| Whisper Small | 17.8% | 3.5% | 0.78 s | On-device option |
| Moonshine V2 Medium | 19.9% | 10.8% | 0.34 s | On-device option |
| Moonshine V2 Small | 26.9% | 12.6% | 0.24 s | On-device option |
| Moonshine V2 Tiny | 45.5% | 25.0% | 0.12 s | On-device option |
- Every model got three to six times worse. Whisper Large V3 Turbo went from 2.2% to 12.9%; Parakeet V3 from 4.4% to 13.6%. A benchmark on clean audio tells you about the benchmark.
- The gap between Whisper and Parakeet closed. On a real voice they are within a point of each other — 12.9% against 13.6% — while Parakeet answers in 0.13 s and Whisper in 2.9 s.
- Gemma's clean-audio lead did not survive. Gemma 4 E4B scored 2.4% on computer voices and 16.1% on a real one, behind both Whisper and Parakeet.
- Apple's built-in recognizer is the fastest thing we measured at 0.10 s per clip, and lands between Parakeet and Gemma on accuracy.
One speaker, one microphone, 20 clips. Errors are counted against the script, so the speaker's own slips count against every model equally — read the ranking, not the absolute numbers. Parakeet V2 and Whisper Large V3 were not run on this set.
Does AI cleanup help, and what does it cost?
A transcript is not yet writing. People say “send it to John, no wait, Sarah” and mean “send it to Sarah”; they say “four thirty PM” and mean “4:30pm”. So we scored the same 20 real-voice clips a second way: against the text the speaker meant, plus 12 pass-or-fail checks of exactly those things. The cleanup step runs on the computer too — two sizes of Google's Gemma 4, and Apple's built-in model for comparison.
| Pipeline | Errors vs intended | Checks passed | Time |
|---|---|---|---|
| Whisper Large V3 Turbo + Gemma 4 E4B | 10.6% | 10/12 | 3.8 s |
| Parakeet V3 + Gemma 4 E2B | 12.1% | 8/12 | 0.56 s |
| Parakeet V3 + Gemma 4 E4B | 12.8% | 7/12 | 0.91 s |
| Parakeet V3, no AI cleanup | 16.1% | 7/12 | 0.13 s |
| Apple SpeechAnalyzer, no AI cleanup | 19.4% | 6/12 | 0.10 s |
| Parakeet V3 + Apple Intelligence | 24.9% | 6/12 | 1.3 s |
- Cleanup is worth about four errors in every hundred words, for roughly half a second: Parakeet alone scored 16.1%, and 12.1% with Gemma 4 E2B behind it.
- The most accurate setup was also the slowest. Whisper Large V3 Turbo with Gemma 4 E4B passed 10 of 12 checks but took 3.8 s — almost all of it Whisper. Swapping in Parakeet gives up a point or two and answers in under a second.
- Not every on-device model can do this job. Apple's built-in model made the text worse than no cleanup at all (24.9%): it dropped words the speaker said, turned a spoken list into numbered points, and refused one ordinary sentence as unsafe.
Twenty clips from one speaker is a small sample, and the two Gemma sizes are within its noise of each other. The differences that are not noise: cleanup against none, and Gemma against Apple's model.
Is cloud dictation faster than on-device?
We assumed it would be. It is not. Walkie logs where the time goes on every Cloud Mode dictation, and across 138 real ones the median was 0.81 s from releasing the key to having the text: 0.24 s transcribing, 0.26 s on AI cleanup, and 0.24 s simply travelling there and back.
On the same laptop, on-device transcription with Parakeet took 0.13 s–0.35 s. Transcribing takes about as long in either place; the cloud then adds the trip there and back. For anything as short as a dictation, a fast local model is hard to beat.
How accurate are on-device models on real speech?
The Hugging Face Open ASR Leaderboard scores 66 speech models on the same real recordings — meetings, earnings calls, audiobooks, podcasts. It is the closest thing the field has to a neutral referee, and it includes the paid cloud engines. The table below is a selection; the full board is linked in the sources.
| Model | Maker | Runs | Errors | Speed | Size | License | In Walkie |
|---|---|---|---|---|---|---|---|
| Scribe v2 Pro | Zoom | Cloud API | 3.59% | — | — | Proprietary | — |
| Azure Speech (07-2026) | Microsoft | Cloud API | 3.81% | — | — | Proprietary | — |
| Scribe v2 | ElevenLabs | Cloud API | 3.97% | — | — | Proprietary | — |
| Qwen3-ASR 1.7B | Alibaba | Runs locally | 4.31% | 820× | 2.04B | Apache 2.0 | — |
| Universal-3.5 Pro | AssemblyAI | Cloud API | 4.34% | — | — | Proprietary | — |
| Canary-Qwen 2.5B | NVIDIA | Runs locally | 4.43% | 867× | 2.5B | CC-BY-4.0 | — |
| Solaria-3 | Gladia | Cloud API | 4.57% | — | — | Proprietary | — |
| Parakeet TDT 0.6B V2 | NVIDIA | Runs locally | 4.70% | 6,025× | 0.6B | CC-BY-4.0 | — |
| Parakeet TDT 0.6B V3 | NVIDIA | Runs locally | 4.86% | 6,076× | 0.6B | CC-BY-4.0 | On-device default |
| Enhanced | Speechmatics | Cloud API | 5.25% | — | — | Proprietary | — |
| Voxtral Mini 3B | Mistral | Runs locally | 5.54% | 181× | 5B | Apache 2.0 | — |
| Whisper Large V3 | OpenAI | Runs locally | 5.78% | 470× | 2B | Apache 2.0 | — |
| Whisper Large V3 Turbo | OpenAI | Runs locally | 6.36% | 797× | 0.8B | MIT | On-device option · Cloud Mode |
| Voxtral Mini 4B Realtime | Mistral | Runs locally | 6.46% | 103× | 4B | Apache 2.0 | — |
| Nemotron 3.5 ASR Streaming | NVIDIA | Runs locally | 7.88% | 1,345× | 0.64B | OpenMDW-1.1 | — |
| Gemma 4 E4B | Runs locally | 9.51% | 161× | 7.94B | Apache 2.0 | — | |
| Gemma 4 E2B | Runs locally | 11.64% | 191× | — | Apache 2.0 | — |
| Model | Maker | Runs | Errors | Speed | Size | License | In Walkie |
|---|---|---|---|---|---|---|---|
| Qwen3-ASR 1.7B | Alibaba | Runs locally | 4.31% | 820× | 2.04B | Apache 2.0 | — |
| Canary-Qwen 2.5B | NVIDIA | Runs locally | 4.43% | 867× | 2.5B | CC-BY-4.0 | — |
| Parakeet TDT 0.6B V2 | NVIDIA | Runs locally | 4.70% | 6,025× | 0.6B | CC-BY-4.0 | — |
| Parakeet TDT 0.6B V3 | NVIDIA | Runs locally | 4.86% | 6,076× | 0.6B | CC-BY-4.0 | On-device default |
| Voxtral Mini 3B | Mistral | Runs locally | 5.54% | 181× | 5B | Apache 2.0 | — |
| Whisper Large V3 | OpenAI | Runs locally | 5.78% | 470× | 2B | Apache 2.0 | — |
| Whisper Large V3 Turbo | OpenAI | Runs locally | 6.36% | 797× | 0.8B | MIT | On-device option · Cloud Mode |
| Voxtral Mini 4B Realtime | Mistral | Runs locally | 6.46% | 103× | 4B | Apache 2.0 | — |
| Nemotron 3.5 ASR Streaming | NVIDIA | Runs locally | 7.88% | 1,345× | 0.64B | OpenMDW-1.1 | — |
| Gemma 4 E4B | Runs locally | 9.51% | 161× | 7.94B | Apache 2.0 | — | |
| Gemma 4 E2B | Runs locally | 11.64% | 191× | — | Apache 2.0 | — |
| Model | Maker | Runs | Errors | Speed | Size | License | In Walkie |
|---|---|---|---|---|---|---|---|
| Scribe v2 Pro | Zoom | Cloud API | 3.59% | — | — | Proprietary | — |
| Azure Speech (07-2026) | Microsoft | Cloud API | 3.81% | — | — | Proprietary | — |
| Scribe v2 | ElevenLabs | Cloud API | 3.97% | — | — | Proprietary | — |
| Universal-3.5 Pro | AssemblyAI | Cloud API | 4.34% | — | — | Proprietary | — |
| Solaria-3 | Gladia | Cloud API | 4.57% | — | — | Proprietary | — |
| Enhanced | Speechmatics | Cloud API | 5.25% | — | — | Proprietary | — |
| Model | Runs | French | German | Spanish | Italian | Portuguese | Dutch | Average | In Walkie |
|---|---|---|---|---|---|---|---|---|---|
| Azure Speech (07-2026) | Cloud API | 2.48% | 1.83% | 1.75% | 0.87% | 2.08% | 2.82% | 1.97% | — |
| ElevenLabs Scribe v2 | Cloud API | 2.93% | 2.30% | 1.85% | 0.90% | 2.60% | 2.51% | 2.18% | — |
| AssemblyAI Universal-3.5 Pro | Cloud API | 2.82% | 2.17% | 2.04% | 1.01% | 2.80% | 3.43% | 2.38% | — |
| Whisper Large V3 | Runs locally | 4.84% | 3.20% | 2.30% | 2.28% | 3.50% | 4.37% | 3.41% | — |
| Voxtral Mini 3B | Runs locally | 4.13% | 3.64% | 3.25% | 2.22% | 3.52% | 5.59% | 3.73% | — |
| Qwen3-ASR 1.7B | Runs locally | 4.06% | 3.35% | 2.92% | 2.23% | 3.97% | 6.59% | 3.85% | — |
| Whisper Large V3 Turbo | Runs locally | 4.90% | 3.67% | 2.73% | 3.30% | 3.77% | 4.96% | 3.89% | On-device option · Cloud Mode |
| Parakeet TDT 0.6B V3 | Runs locally | 4.68% | 4.16% | 3.25% | 2.42% | 4.46% | 6.39% | 4.23% | On-device default |
| Gemma 4 E4B | Runs locally | 7.42% | 5.75% | 3.64% | 3.35% | 4.52% | 8.99% | 5.61% | — |
| Nemotron 3.5 ASR Streaming | Runs locally | 9.86% | 8.46% | 4.23% | 4.81% | 5.80% | 11.90% | 7.51% | — |
What the leaderboard says
- The cloud engines lead. Zoom, Microsoft and ElevenLabs score between 3.59% and 3.97% — better than anything you can run yourself. The cost is that your audio lives, however briefly, on their servers.
- Parakeet is the most accurate model in its speed class. Parakeet V2 and V3 score 4.7% and 4.86% while processing audio about 6,025 times faster than real time on the benchmark's GPU. Every commercially usable model that is more accurate than Parakeet V2 is at least 2.9 times slower; the most accurate open model, Alibaba's Qwen3-ASR 1.7B (4.31%), is about 7 times slower and three times larger. Parakeet V3 ranks 9 of the 17 shown here, ahead of every Whisper model.
- Whisper Large V3 wins on European languages. It averages 3.41% across six languages — the best of any open model — against 3.89% for Turbo and 4.23% for Parakeet V3. Walkie offers Turbo rather than Large: on our laptop it was faster and at least as accurate, and on this set it trails Large by about half a point.
- Clean-audio results do not transfer. Gemma 4 E4B, which scored 2.4% on our computer voices, scores 9.51% on real speech — about twice Parakeet's error rate.
Why not use a chatbot model like Gemma or Voxtral?
Models such as Google's Gemma 4 and Mistral's Voxtral can listen as well as read, which makes it tempting to use one model for everything. For dictation, three things get in the way:
- Accuracy on real speech. On the leaderboard, Gemma 4 E4B scores 9.51% and E2B 11.64%, against 4.86% for Parakeet V3. Voxtral Mini 3B does better at 5.54%, but runs at 181× against Parakeet's 6,076×.
- A 30-second limit. Google documents a maximum of 30 seconds of audio per request for Gemma. Long dictations have to be cut up, and every cut is a place to lose a word.
- Size and speed. Gemma 4 E4B is a 5.2 GB download and took 1.3 s per clip on our laptop. Parakeet V3 is 478 MB and took 0.16 s.
Language models are excellent at a different job: understanding and rewriting text. The strongest setup pairs a dedicated speech model that listens with a language model that writes — which is exactly how Walkie's Cloud Mode works.
What about Nemotron 3.5 and Moonshine?
NVIDIA Nemotron 3.5 is a streaming model: words appear while you are still talking, in up to 40 languages. It scores 7.88% on English and 7.51% on the European set — behind Parakeet on both. Walkie shows your text when you finish speaking, and Parakeet finishes a whole clip in 0.16 s, so streaming would not make dictation feel faster.
Moonshine is tiny and quick. It stays in Walkie for US English on older computers, where a 31 MB download matters more than accents it was not trained on.
Cloud Mode: what runs, and why
Cloud Mode transcribes with Whisper Large V3 Turbo on Groq's hardware and switches to full Whisper Large V3 when it suspects Turbo has translated instead of transcribed. An AI cleanup step then adds punctuation, removes filler words and formats the text for the app you are typing into. Cloud requests run under zero data retention, and nothing is used to train AI models.
Whisper Turbo is not the most accurate engine on the leaderboard — it scores 6.36% to Azure's 3.81%. We use it because it is fast, covers 99 languages, and the cleanup step does what no raw transcript does: turns speech into writing. If raw English accuracy is what you are after, Parakeet in On-Device Mode scores better than Whisper Turbo on the same benchmark — and never leaves your computer.
How Walkie compares to Wispr Flow and others
| App | Where it transcribes | Choose your model | Speech engine |
|---|---|---|---|
| Walkie | Your computer or the cloud — your choice | Yes, 7 on-device models | Parakeet, Whisper, Moonshine; Whisper Turbo in Cloud Mode |
| Wispr Flow | Cloud only | No | One proprietary cloud model, not disclosed |
| Apple Dictation | On your device, on modern Apple hardware | No | Apple's own model |
| Handy | Your computer | Yes | Whisper or Parakeet |
| Superwhisper, MacWhisper, VoiceInk | Your computer | Varies by app | Whisper-based; other options vary |
The honest version: a cloud-only tool like Wispr Flow may well run an engine as accurate as the leaders above — it just does not say, and every word you dictate goes to its servers to find out. Apple Dictation is free and already installed. What Walkie adds is the choice: the fastest accurate local model by default, Whisper when you need 99 languages or rare words, the cloud when you want polish — and the numbers to decide with.
For a feature-by-feature look at the apps that run these models on your computer — Handy, VoiceInk, MacWhisper, Superwhisper and more — see the best local dictation apps, compared.
Which model should you pick?
- You dictate in English or another of 25 European languages and want it instant: Parakeet V3, Walkie's default.
- Other languages, or lots of names and technical terms: Whisper Small if you want speed, Whisper Large V3 Turbo if you want accuracy — and add the words you use to your custom dictionary.
- An older or low-memory computer: Whisper Small, or Moonshine for US English.
- The best-written result, and the cloud is fine: Cloud Mode.
Method and sources
Laptop test, 2026-09-19: 80 clips (2.4–8.2 seconds each) generated with the Samantha (US), Daniel (UK), Rishi (India), Karen (Australia) voices, 16 kHz mono. Every model used the exact file Walkie downloads, including Walkie's quantized Whisper Medium and Large, which is why our Whisper numbers are Walkie's, not the full-precision models'. Each model was loaded once, warmed with one clip, then timed per clip; the table shows the median. Gemma ran through llama.cpp with Google's recommended transcription prompt. Scoring used the Whisper English normalizer — the one the leaderboard uses — and word error rate across all clips.
Leaderboard figures were retrieved 2026-09-19 from the published English results and multilingual results files. Soniox and Deepgram are not on the leaderboard, so they are not in these tables.
The 20 test phrases
- um so I think we need to uh update the docs before Friday
- So basically the plan is to ship the new version on Friday.
- Send it to John, no wait, send it to Sarah.
- It costs fifty, no, sixty dollars.
- Let's meet Monday, actually Tuesday, at 4:30 PM.
- I like the blue one better than the green one.
- Turn right at the light, then go two blocks.
- Summarize this for me and tell me what the weather is.
- Send it to example.com, and the budget is $25,000.
- Hey John, just wanted to say thanks for the call today. Talk soon, Adam.
- Hey Sarah, can you move our Thursday sync to 3 PM and send the updated deck to the whole team before then?
- Quarterly revenue grew 18 percent to 2.4 million dollars, mostly from the enterprise tier.
- Please add Kubernetes, PostgreSQL, and the Redis cluster to the migration checklist for next sprint.
- My flight lands at 7:45 on Friday, so I'll probably miss the first half of the offsite.
- Remind me to call Dr. Patel about the lab results and to renew the prescription for amoxicillin.
- I think the second option is better because it's cheaper, faster to build, and easier to maintain over time.
- Can you summarize the three main risks in the proposal and suggest one mitigation for each?
- Order two large pizzas, one margherita and one pepperoni, and have them delivered to 1200 Market Street.
- We need to finish the onboarding redesign, fix the login bug on Android, and ship version 2.1 by the end of the month.
- Thanks for the thoughtful feedback on the draft. I've incorporated most of it, and I'll send a new version tomorrow morning.
Questions
On the Open ASR Leaderboard's English speech, the most accurate open model is Alibaba's Qwen3-ASR 1.7B (4.31% word error rate), with several 2–5 billion-parameter models close behind, such as NVIDIA's Canary-Qwen 2.5B (4.43%). NVIDIA's Parakeet TDT 0.6B V2 scores 4.7% while running about 7 times faster than Qwen3-ASR. For European languages, Whisper Large V3 is the most accurate open model at 3.41% averaged across French, German, Spanish, Italian, Portuguese and Dutch.
For fast dictation, usually yes. On one ordinary laptop Parakeet V3 transcribed a typical clip in 0.16 s, 19 times faster than Whisper Large V3 Turbo, and on the leaderboard's real English speech Parakeet makes fewer errors (4.86% against 6.36% for Whisper Turbo). Whisper is the better choice for rare words and technical jargon, for languages outside Parakeet's 25, and for European languages when accuracy matters more than speed.
Yes. Gemma 4 E2B and E4B accept audio directly, up to 30 seconds per request. On clean test clips Gemma 4 E4B was about as accurate as Whisper Large, but on the leaderboard's real-world English speech it scores 9.51%, roughly twice Parakeet's error rate. It is a better fit for understanding and rewriting text than for word-for-word transcription.
Wispr Flow runs a single proprietary cloud model and does not say which one. Every dictation goes to its servers, and there is no model to choose between.
Yes. In On-Device Mode the speech model runs on your own computer and makes no network request, so it works with the wifi off. On-Device Mode dictation is free forever, with no account needed.
Not in our measurements. Across 138 real Cloud Mode dictations the median was 0.81 s from releasing the key to having the text, of which 0.24 s was the network round trip. On the same laptop, Parakeet transcribed on-device in 0.13 s to 0.35 s.
On a base Apple M4 laptop, Parakeet V3 transcribed clips of 2.4 to 8.2 seconds in a median 0.16 s, about 33 times faster than real time.