A5S V2: state-of-the-art streaming STT on consumer devices
Evaluation on noisy, far-field, and multi-speaker English speech, with inference measured on an M4 MacBook Air and an iPhone 16.
A5S V2 and proprietary cloud streaming systems
Coval comparison as of September 1, 2026, using the August 29 leaderboard snapshot.
- Task
- Streaming English Automatic Speech Recognition (ASR)
- Primary metric
- Word error rate (WER)
- Evaluation set
- 133,027 words from 4 public STT corpora
00
Abstract
A5S V2 is a streaming automatic speech recognition model developed for speech in difficult real-world environments: distant microphones, room reverberation, background noise, and competing speakers. The model is small enough to run locally on current consumer devices such as Apple iPhones or MacBooks. This note reports its accuracy on an evaluation set composed of data from four public English speech-recognition corpora and its latency on the Coval Speech-to-Text benchmark.
On an evaluation set built from 4 open-source audio corpora representing real-world, noisy speech, and multi-talker meetings (Mega-ASR, AMI, DiPCo, and NOTSOFAR), A5S V2 obtains 26.27% word error rate (WER). Despite running on consumer edge hardware, A5S V2 also outperforms all eight evaluated SOTA proprietary cloud streaming systems, with a 0.37 percentage-point lead over ElevenLabs Scribe v2 Realtime. It ranks first in the open and open-weight comparison, with a 36.7% relative WER reduction versus the second-best system.
On an M4 MacBook Air, A5S V2's median Time to Final Segment (TTFS) is 45.50 ms across 947 clips. On a physical iPhone 16, the corresponding result is 50.01 ms. This is faster than every model benchmarked by Coval on the same dataset (as of September 1, 2026), including models run on GPUs in the cloud.
We publish the fixed references, raw predictions, provider request configurations, timing records, checksums, and scoring code. For select enterprises and startups, A5S V2 is available for evaluation through the AirCaps Playground dashboard. Developers should contact AirCaps to request access. Deployment on supported edge hardware is also available on request.
01
Motivation
1. Real-world performance gap
Speech recognition is easiest when there is one clear speaker, the microphone is close by, and the recording is clean. Those conditions describe audiobooks, podcasts, and call-center audio. They could not be further from the reality of audio captured by microphones in real meeting rooms, cars, consumer appliances, robots, or wearable devices. In a room, the direct voice may be weaker than its reflections. Noise can be diffuse or directional. Other speakers occupy the same frequency range as the target voice. Distance removes high-frequency detail and lowers the signal-to-noise ratio. These effects occur together, and they vary while a person or device moves.
While state-of-the-art STT models report benchmark-breaking numbers on legacy, clean datasets, they completely break down in "hard" acoustical situations. For example, OpenAI's latest streaming STT model, GPT Live Transcribe (released Aug 2026), achieves 2.6% WER on VoxPopuli according to Artificial Analysis. Here's a sample from VoxPopuli:
Loading waveform…
The same model achieves 40.92% WER (15x worse) on a subset of DiPCo (Dinner Party Corpus), which sounds a bit more like this:
Loading waveform…
2. Edge deployment
Cloud models have 3 disadvantages: (1) they require consistent internet connections (WiFi or cellular), (2) they are expensive to use on a per-hour basis, and (3) they are not privacy-preserving.
However, local models have generally introduced another accuracy loss. The strongest speech systems are commonly closed, multi-billion-parameter proprietary cloud services charged by usage. Historically, edge-deployable models had to be tiny. Older edge models like Vosk and Kaldi were leagues behind cloud models (WERs on our eval set would have been close to 100%).
OpenAI's Whisper was the first competitive open-source model, but the Large variant could not realistically be deployed on consumer hardware, and the model was not meant for real-time streaming (smaller variants were much less accurate). Nvidia's Parakeet V3 launched in August 2025 was the first glimpse of consumer-hardware-deployable STT that could give cloud models a run for their money, but it was an asynchronous (batch/offline) model.
3. Streaming inference
Streaming adds a separate constraint. At time t, a causal model can use only the audio received at or before t. It cannot inspect the remainder of an utterance before deciding what the current words are. Batch-model accuracy has therefore always held a 10-40% relative WER advantage over streaming models.
Our goal with A5S V2 was to test whether it was possible to build a model that performs well on all three fronts: real-world accuracy, edge deployment on consumer hardware, and streaming inference.
02
Accuracy evaluation
Against eight leading proprietary cloud services, A5S V2 ranks first and outperforms all eight systems. It records 1.4% lower WER than the next-ranked cloud system, ElevenLabs Scribe v2 Realtime, a 0.37-point absolute difference.
A5S V2 records 26.27% WER and ranks first ahead of four open and open-weight streaming systems evaluated. The next-best system, NVIDIA Nemotron 3 ASR Streaming 0.6B, records 41.52% WER: a 15.25-point absolute difference and a 36.7% relative reduction in WER.
Proprietary cloud systems
WER (%) · lower is better
Open and open-weight systems
WER (%) · lower is better
Evaluation data
The primary evaluation uses four public STT corpora with approximately equal normalized word counts. Mega-ASR contributes short utterances under five acoustic conditions. AMI, DiPCo, and NOTSOFAR contribute complete meetings. Meeting audio is not segmented or cleaned before inference.
| Corpus | Selection | Audio used | Reference words |
|---|---|---|---|
| Mega-ASR | 1,250 utterances; 250 per acoustic condition | Original mono recordings | 32,928 |
| AMI | 7 unseen-evaluation scenario meetings | Single distant microphone, Array1-01 | 32,928 |
| DiPCo | 5 evaluation meetings and 1 development meeting | U01.CH1 with fixed +5.1 dB gain | 33,679 |
| NOTSOFAR | 22 eval-small meetings | Single conference-device channel | 33,492 |
Table 1. Fixed corpus selection. Each meeting is streamed as one complete recording. No denoising, beamforming, silence removal, or reference prompting is used.
Selection of comparison systems
We deliberately selected leading streaming ASR systems rather than sampling available models at random. Accuracy was the primary criterion: we sought the newest, strongest proprietary cloud services and open or open-weight models with credible realtime performance, using independent accuracy and latency benchmarks to identify systems at or near the state of the art.
Proprietary cloud
Six of the eight proprietary cloud systems were in the top 10 of the Artificial Analysis AA-WER Streaming Index as of September 5, 2026: Meta Muse ranked 1st, ElevenLabs 3rd, OpenAI 5th, xAI Grok 6th, Gemini 7th, and AssemblyAI 8th. Google Chirp ranked 12th and Deepgram Nova-3 ranked 19th.
| System in this evaluation | AA-WER rank | AA-WER | Coval WER rank | Coval TTFS rank | Pipecat WER rank | Pipecat TTFS rank |
|---|---|---|---|---|---|---|
| Meta Muse Transcribe | 1 | 3.1% | — | — | — | — |
| ElevenLabs Scribe v2 Realtime | 3 | 3.6% | 15 | 7 | 14 | 6 |
| OpenAI GPT Live Transcribe | 5 | 3.9% | — | — | — | — |
| xAI Grok Speech-to-Text Streaming | 6 | 3.9% | — | — | — | — |
| Google Gemini 3.5 Transcribe Live | 7 | 4.0% | — | — | — | — |
| AssemblyAI Universal-3.5 Pro Realtime | 8 | 4.0% | 1 | 8 | 2 | 7 |
| Google Chirp 3 Streaming | 12 | 4.8% | 3 | 24 | — | — |
| Deepgram Nova-3 Streaming | 19 | 6.6% | 18 | 4 | 5 | 3 |
AA-WER ranks are from September 5, 2026; Coval and Pipecat ranks retain their August 2026 snapshots. Dashes indicate that the exact endpoint was not ranked. External benchmarks use different data, normalization, settings, and model versions, so these values explain system selection rather than provide directly comparable scores.
Open source and open weight
| System in this evaluation | AA-WER rank | Coval WER rank | Coval TTFS rank | Pipecat WER rank | Pipecat TTFS rank |
|---|---|---|---|---|---|
| NVIDIA Nemotron 3 ASR Streaming 0.6B | 1 of 2 | — | — | 1 of 2 | 1 of 2 |
| Whisper Large V3 with Whisper-Streaming | — | 2 of 2* | 1 of 2* | — | — |
| Mistral Voxtral Mini Transcribe Realtime 2602 | 2 of 2 | 1 of 2 | 2 of 2 | 2 of 2 | 2 of 2 |
| Kyutai STT 2.6B English | — | — | — | — | — |
Ranks are recalculated only within the evaluated open/open-weight model families present in each external benchmark: Artificial Analysis (Nemotron and Mistral), Coval (Mistral and Whisper Large V3), and Pipecat (Nemotron and Mistral). Dashes indicate no matching model. *Coval serves Whisper Large V3 through Together AI rather than the Whisper-Streaming implementation used here.
Proprietary cloud system results
Each service used its documented streaming interface and the same audio, pacing, and text normalization. One saved run is reported; every system completed every evaluation item.
| System | Mega-ASR | AMI | DiPCo | NOTSOFAR | WER |
|---|---|---|---|---|---|
| A5S V2 | 19.65 | 20.31 | 33.37 | 31.75 | 26.27 |
| ElevenLabs Scribe v2 Realtime | 22.64 | 20.41 | 31.71 | 31.81 | 26.64 |
| Meta Muse Transcribe | 25.38 | 26.06 | 32.71 | 36.05 | 30.05 |
| AssemblyAI Universal-3.5 Pro Realtime | 19.25 | 29.83 | 36.61 | 36.23 | 30.48 |
| OpenAI GPT Live Transcribe | 27.81 | 30.16 | 39.86 | 40.85 | 34.67 |
| xAI Grok Speech-to-Text Streaming | 26.37 | 38.24 | 64.78 | 37.95 | 41.84 |
| Deepgram Nova-3 Streaming | 40.46 | 35.95 | 68.19 | 37.68 | 45.57 |
| Google Chirp 3 Streaming | 34.68 | 33.93 | 79.19 | 44.85 | 48.16 |
| Google Gemini 3.5 Transcribe Live | 35.32 | 60.16 | 96.28 | 63.07 | 63.71 |
Table 2. Pooled WER (%) within each corpus; the final column is the unweighted mean of the four corpus values. Lower is better.
Comparison Samples
These examples compare A5S V2 with four leading cloud models on difficult real-world audio, highlighting cases where A5S V2 transcribes speech correctly while other systems miss or misrecognize it. These samples deliberately include background music or distractions and degraded / far field speech.
Open and open-weight system results
Nemotron and Voxtral were evaluated through provider-operated endpoints serving the identified open-weight models. Whisper and Kyutai were run locally using the cited public implementations. Every system processed the same 133,027 normalized reference words.
| System | Mega-ASR | AMI | DiPCo | NOTSOFAR | WER |
|---|---|---|---|---|---|
| A5S V2 | 19.65 | 20.31 | 33.37 | 31.75 | 26.27 |
| NVIDIA Nemotron 3 ASR Streaming 0.6B | 24.49 | 33.32 | 61.72 | 46.58 | 41.52 |
| Whisper Large V3 with Whisper-Streaming | 28.06 | 55.86 | 75.69 | 42.05 | 50.42 |
| Mistral Voxtral Mini Transcribe Realtime 2602 | 32.55 | 48.71 | 75.52 | 45.30 | 50.52 |
| Kyutai STT 2.6B English | 33.18 | 38.54 | 70.36 | 63.36 | 51.36 |
Table 3. Pooled WER (%) within each corpus; the final column is the unweighted mean of the four corpus values. Lower is better.
03
On-device latency
A5S V2 records 43.99 ms mean TTFS on an M4 MacBook Air and 50.39 ms on an iPhone 16. Inserted into Coval's August 29, 2026 ranking, the two device trials place first and second, ahead of all 24 listed systems (comparison as of September 1, 2026). The Mac result is 31.3% lower than the 64 ms leader; the iPhone result is 21.3% lower.
Both measurements are fully on device. There is no cloud inference or network round trip. Each consumer device processed all 947 items locally without a failed item. ElevenLabs Scribe v2 Realtime requires 2.7× as much finalization time as the Mac and 2.4× as much as the iPhone.
A5S V2 inserted into the Coval TTFS ranking
Mean TTFS · lower is better
- 1
A5S V2 · MacBook Air M4 device trial
43.99 - 2
A5S V2 · iPhone 16 device trial
50.39 - 3
Soniox STT RT v5
64 - 4
NVIDIA Parakeet TDT 0.6B v3
70 - 5
Inworld STT 1
83 - 6
Deepgram Nova 3
99 - 7
Deepgram Nova 2
101 - 8
Cartesia Ink 2
108 - 9
ElevenLabs Scribe v2 Realtime
120 - 10
AssemblyAI Universal 3.5 Pro
146 ms
Measurement protocol
TTFS follows Coval's definition: elapsed time from the reference end of speech to receipt of the final transcription segment. Each trial used 100 ms real-time audio pacing and concurrency one. Reference endpoint annotations isolate model finalization and exclude live endpoint detection, model loading, and warm-up.
| Device | Items | TTFS mean | TTFS p50 | TTFS p95 |
|---|---|---|---|---|
| MacBook Air, M4 | 947 | 43.99 | 45.50 | 53.28 |
| iPhone 16, A18 | 947 | 50.39 | 50.01 | 57.23 |
Table 4. TTFS in milliseconds, item-weighted across Coval stt-v1 and stt-v3. Lower is better.
MacBook Air · M4
947-item Coval trial
iPhone 16 · A18
947-item Coval trial
The live recordings use a 250 ms silence threshold to identify the end of speech. Their displayed values are implementation demonstrations and are distinct from the oracle-endpoint measurements in Table 4.
Pareto frontier: latency vs accuracy (Coval)
Lower-left is better
04
Deployment properties
A locally executed speech model avoids an audio upload and continues to operate without a network connection. It also removes the provider's per-minute inference charge. These properties are separate from accuracy and are not represented in WER.
The proprietary cloud systems with public list prices publish rates ranging from $0.20 to $1.02 per audio hour under the plans shown below. A5S V2 has no external API charge per processed minute, although local compute, energy, integration, and hardware costs remain.
| System used in evaluation | Deployment in this comparison | Published usage price |
|---|---|---|
| A5S V2 | Local, Mac or iPhone | No external per-minute charge |
| xAI Grok Speech-to-Text Streaming | Proprietary cloud API | $0.20 / audio hour |
| Deepgram Nova-3 | Proprietary cloud API | $0.29 / audio hour |
| ElevenLabs Scribe v2 Realtime | Proprietary cloud API | $0.39 / audio hour |
| AssemblyAI Universal-3.5 Pro | Proprietary cloud API | $0.45 / audio hour |
| Google Gemini 3.5 Transcribe Live | Proprietary cloud API | ~$0.54 / audio hour |
| Google Chirp 3 | Proprietary cloud API | $0.96 / audio hour |
| OpenAI GPT Live Transcribe | Proprietary cloud API | $1.02 / audio hour |
| Meta Muse Transcribe | Proprietary cloud API | Not publicly listed |
Table 5. Public list or pay-as-you-go prices checked in September 2026. Gemini's blended rate is the provider's token-based estimate. Discounts, add-ons, and enterprise terms vary.
05
Limitations
Release v1 reports one saved inference trial for each system and corpus. Confidence intervals are not included. The planned analysis uses paired cluster bootstrap resampling by utterance for Mega-ASR and by complete meeting for the other corpora.
The Mega-ASR subset was sampled from a public training split because the standard test set did not separate robust systems in our preliminary experiments. A5S V2 was not trained, fine-tuned, selected, or prompted on the 1,250 chosen recordings or transcripts. We cannot determine whether third-party systems encountered these items or related upstream data. Mega-ASR should be read as a developer-created robustness diagnostic alongside the three meeting corpora, not as a speaker-disjoint test set.
These results measure English streaming transcription without diarization. They do not establish performance for other languages, speaker attribution, live endpoint detection, or downstream semantic tasks. Results are specific to the fixed data, normalization, provider versions, and request configurations in the release.
07
Sources
- AirCaps. A5S V2 ASR benchmark: code, protocol, provider runners, and scores.
- AirCaps. A5S V2 ASR benchmark dataset: audio, references, predictions, latency trials, and checksums.
- Coval. Voice AI benchmark metrics and current speech-to-text results.
- Artificial Analysis. AA-WER Streaming Index and time-to-final-transcription comparison.
- Pipecat. Streaming STT benchmark on the smart-turn voice-agent dataset.
- Meta. Muse Transcribe streaming speech-to-text documentation.
- xAI. Streaming speech-to-text documentation.
- Google. Gemini Live transcription documentation.
- Xie et al. Mega-ASR: Towards In-the-wild² Speech Recognition via Scaling up Real-world Acoustic Simulation.
- Carletta et al. The AMI Meeting Corpus.
- Dinner Party Corpus: a multi-view multi-listener speech corpus.
- NOTSOFAR-1 Challenge: new datasets, baseline, and tasks for distant meeting transcription.