Research note

September 2026

A5S V2: state-of-the-art streaming STT on consumer devices

Evaluation on noisy, far-field, and multi-speaker English speech, with inference measured on an M4 MacBook Air and an iPhone 16.

47.9%Relative WER reduction vs. OpenAI Whisper Large V3
45.50 msMedian TTFS on M4 MacBook Air
AccuracyWER (%) · lower is better

A5S V2 and proprietary cloud streaming systems

Finalization latencyMean TTFS (ms) · lower is better

Coval comparison as of September 1, 2026, using the August 29 leaderboard snapshot.

Task
Streaming English Automatic Speech Recognition (ASR)
Primary metric
Word error rate (WER)
Evaluation set
133,027 words from 4 public STT corpora

A5S V2 is a streaming automatic speech recognition model developed for speech in difficult real-world environments: distant microphones, room reverberation, background noise, and competing speakers. The model is small enough to run locally on current consumer devices such as Apple iPhones or MacBooks. This note reports its accuracy on an evaluation set composed of data from four public English speech-recognition corpora and its latency on the Coval Speech-to-Text benchmark.

On an evaluation set built from 4 open-source audio corpora representing real-world, noisy speech, and multi-talker meetings (Mega-ASR, AMI, DiPCo, and NOTSOFAR), A5S V2 obtains 26.27% word error rate (WER). Despite running on consumer edge hardware, A5S V2 also outperforms all eight evaluated SOTA proprietary cloud streaming systems, with a 0.37 percentage-point lead over ElevenLabs Scribe v2 Realtime. It ranks first in the open and open-weight comparison, with a 36.7% relative WER reduction versus the second-best system.

On an M4 MacBook Air, A5S V2's median Time to Final Segment (TTFS) is 45.50 ms across 947 clips. On a physical iPhone 16, the corresponding result is 50.01 ms. This is faster than every model benchmarked by Coval on the same dataset (as of September 1, 2026), including models run on GPUs in the cloud.

We publish the fixed references, raw predictions, provider request configurations, timing records, checksums, and scoring code. For select enterprises and startups, A5S V2 is available for evaluation through the AirCaps Playground dashboard. Developers should contact AirCaps to request access. Deployment on supported edge hardware is also available on request.

1. Real-world performance gap

Speech recognition is easiest when there is one clear speaker, the microphone is close by, and the recording is clean. Those conditions describe audiobooks, podcasts, and call-center audio. They could not be further from the reality of audio captured by microphones in real meeting rooms, cars, consumer appliances, robots, or wearable devices. In a room, the direct voice may be weaker than its reflections. Noise can be diffuse or directional. Other speakers occupy the same frequency range as the target voice. Distance removes high-frequency detail and lowers the signal-to-noise ratio. These effects occur together, and they vary while a person or device moves.

While state-of-the-art STT models report benchmark-breaking numbers on legacy, clean datasets, they completely break down in "hard" acoustical situations. For example, OpenAI's latest streaming STT model, GPT Live Transcribe (released Aug 2026), achieves 2.6% WER on VoxPopuli according to Artificial Analysis. Here's a sample from VoxPopuli:

Loading waveform…

0:00 / 0:00
VoxPopuli, English test set. Clean, segmented, single-speaker European Parliament speech. Reference: “Then we, as a Parliament, could take our responsibility and quickly vote through these measures…”

The same model achieves 40.92% WER (15x worse) on a subset of DiPCo (Dinner Party Corpus), which sounds a bit more like this:

Loading waveform…

0:00 / 0:00
DiPCo, recording S01, microphone U01.CH1. Far-field dinner-party speech with room reverberation and several short speaker turns. The released full-meeting GPT Live prediction omits the greeting and brief exchanges in this interval.

2. Edge deployment

Cloud models have 3 disadvantages: (1) they require consistent internet connections (WiFi or cellular), (2) they are expensive to use on a per-hour basis, and (3) they are not privacy-preserving.

However, local models have generally introduced another accuracy loss. The strongest speech systems are commonly closed, multi-billion-parameter proprietary cloud services charged by usage. Historically, edge-deployable models had to be tiny. Older edge models like Vosk and Kaldi were leagues behind cloud models (WERs on our eval set would have been close to 100%).

OpenAI's Whisper was the first competitive open-source model, but the Large variant could not realistically be deployed on consumer hardware, and the model was not meant for real-time streaming (smaller variants were much less accurate). Nvidia's Parakeet V3 launched in August 2025 was the first glimpse of consumer-hardware-deployable STT that could give cloud models a run for their money, but it was an asynchronous (batch/offline) model.

3. Streaming inference

Streaming adds a separate constraint. At time t, a causal model can use only the audio received at or before t. It cannot inspect the remainder of an utterance before deciding what the current words are. Batch-model accuracy has therefore always held a 10-40% relative WER advantage over streaming models.

Our goal with A5S V2 was to test whether it was possible to build a model that performs well on all three fronts: real-world accuracy, edge deployment on consumer hardware, and streaming inference.

Against eight leading proprietary cloud services, A5S V2 ranks first and outperforms all eight systems. It records 1.4% lower WER than the next-ranked cloud system, ElevenLabs Scribe v2 Realtime, a 0.37-point absolute difference.

A5S V2 records 26.27% WER and ranks first ahead of four open and open-weight streaming systems evaluated. The next-best system, NVIDIA Nemotron 3 ASR Streaming 0.6B, records 41.52% WER: a 15.25-point absolute difference and a 36.7% relative reduction in WER.

Proprietary cloud systems

WER (%) · lower is better

Figure 1. WER on the fixed four-corpus evaluation. All systems completed all 1,250 utterances and 35 complete meetings.

Open and open-weight systems

WER (%) · lower is better

Figure 2. WER on the fixed four-corpus evaluation. Values are the unweighted mean of the corpus-level pooled WERs.

Evaluation data

The primary evaluation uses four public STT corpora with approximately equal normalized word counts. Mega-ASR contributes short utterances under five acoustic conditions. AMI, DiPCo, and NOTSOFAR contribute complete meetings. Meeting audio is not segmented or cleaned before inference.

CorpusSelectionAudio usedReference words
Mega-ASR1,250 utterances; 250 per acoustic conditionOriginal mono recordings32,928
AMI7 unseen-evaluation scenario meetingsSingle distant microphone, Array1-0132,928
DiPCo5 evaluation meetings and 1 development meetingU01.CH1 with fixed +5.1 dB gain33,679
NOTSOFAR22 eval-small meetingsSingle conference-device channel33,492

Table 1. Fixed corpus selection. Each meeting is streamed as one complete recording. No denoising, beamforming, silence removal, or reference prompting is used.

Selection of comparison systems

We deliberately selected leading streaming ASR systems rather than sampling available models at random. Accuracy was the primary criterion: we sought the newest, strongest proprietary cloud services and open or open-weight models with credible realtime performance, using independent accuracy and latency benchmarks to identify systems at or near the state of the art.

Proprietary cloud

Six of the eight proprietary cloud systems were in the top 10 of the Artificial Analysis AA-WER Streaming Index as of September 5, 2026: Meta Muse ranked 1st, ElevenLabs 3rd, OpenAI 5th, xAI Grok 6th, Gemini 7th, and AssemblyAI 8th. Google Chirp ranked 12th and Deepgram Nova-3 ranked 19th.

Artificial Analysis AA-WER Streaming Index with all eight evaluated proprietary cloud systems outlined
Figure 3. Artificial Analysis AA-WER Streaming Index with all eight evaluated proprietary cloud systems outlined; six rank in the top 10. Rankings cited as of September 5, 2026. Open the current chart ↗
System in this evaluationAA-WER rankAA-WERCoval WER rankCoval TTFS rankPipecat WER rankPipecat TTFS rank
Meta Muse Transcribe13.1%————
ElevenLabs Scribe v2 Realtime33.6%157146
OpenAI GPT Live Transcribe53.9%————
xAI Grok Speech-to-Text Streaming63.9%————
Google Gemini 3.5 Transcribe Live74.0%————
AssemblyAI Universal-3.5 Pro Realtime84.0%1827
Google Chirp 3 Streaming124.8%324——
Deepgram Nova-3 Streaming196.6%18453

AA-WER ranks are from September 5, 2026; Coval and Pipecat ranks retain their August 2026 snapshots. Dashes indicate that the exact endpoint was not ranked. External benchmarks use different data, normalization, settings, and model versions, so these values explain system selection rather than provide directly comparable scores.

Open source and open weight

System in this evaluationAA-WER rankCoval WER rankCoval TTFS rankPipecat WER rankPipecat TTFS rank
NVIDIA Nemotron 3 ASR Streaming 0.6B1 of 2——1 of 21 of 2
Whisper Large V3 with Whisper-Streaming—2 of 2*1 of 2*——
Mistral Voxtral Mini Transcribe Realtime 26022 of 21 of 22 of 22 of 22 of 2
Kyutai STT 2.6B English—————

Ranks are recalculated only within the evaluated open/open-weight model families present in each external benchmark: Artificial Analysis (Nemotron and Mistral), Coval (Mistral and Whisper Large V3), and Pipecat (Nemotron and Mistral). Dashes indicate no matching model. *Coval serves Whisper Large V3 through Together AI rather than the Whisper-Streaming implementation used here.

Proprietary cloud system results

Each service used its documented streaming interface and the same audio, pacing, and text normalization. One saved run is reported; every system completed every evaluation item.

SystemMega-ASRAMIDiPCoNOTSOFARWER
A5S V219.6520.3133.3731.7526.27
ElevenLabs Scribe v2 Realtime22.6420.4131.7131.8126.64
Meta Muse Transcribe25.3826.0632.7136.0530.05
AssemblyAI Universal-3.5 Pro Realtime19.2529.8336.6136.2330.48
OpenAI GPT Live Transcribe27.8130.1639.8640.8534.67
xAI Grok Speech-to-Text Streaming26.3738.2464.7837.9541.84
Deepgram Nova-3 Streaming40.4635.9568.1937.6845.57
Google Chirp 3 Streaming34.6833.9379.1944.8548.16
Google Gemini 3.5 Transcribe Live35.3260.1696.2863.0763.71

Table 2. Pooled WER (%) within each corpus; the final column is the unweighted mean of the four corpus values. Lower is better.

Comparison Samples

These examples compare A5S V2 with four leading cloud models on difficult real-world audio, highlighting cases where A5S V2 transcribes speech correctly while other systems miss or misrecognize it. These samples deliberately include background music or distractions and degraded / far field speech.

Open and open-weight system results

Nemotron and Voxtral were evaluated through provider-operated endpoints serving the identified open-weight models. Whisper and Kyutai were run locally using the cited public implementations. Every system processed the same 133,027 normalized reference words.

SystemMega-ASRAMIDiPCoNOTSOFARWER
A5S V219.6520.3133.3731.7526.27
NVIDIA Nemotron 3 ASR Streaming 0.6B24.4933.3261.7246.5841.52
Whisper Large V3 with Whisper-Streaming28.0655.8675.6942.0550.42
Mistral Voxtral Mini Transcribe Realtime 260232.5548.7175.5245.3050.52
Kyutai STT 2.6B English33.1838.5470.3663.3651.36

Table 3. Pooled WER (%) within each corpus; the final column is the unweighted mean of the four corpus values. Lower is better.

A5S V2 records 43.99 ms mean TTFS on an M4 MacBook Air and 50.39 ms on an iPhone 16. Inserted into Coval's August 29, 2026 ranking, the two device trials place first and second, ahead of all 24 listed systems (comparison as of September 1, 2026). The Mac result is 31.3% lower than the 64 ms leader; the iPhone result is 21.3% lower.

Both measurements are fully on device. There is no cloud inference or network round trip. Each consumer device processed all 947 items locally without a failed item. ElevenLabs Scribe v2 Realtime requires 2.7× as much finalization time as the Mac and 2.4× as much as the iPhone.

A5S V2 inserted into the Coval TTFS ranking

Mean TTFS · lower is better

  1. 1

    A5S V2 · MacBook Air M4 device trial

    43.99
  2. 2

    A5S V2 · iPhone 16 device trial

    50.39
  3. 3

    Soniox STT RT v5

    64
  4. 4

    NVIDIA Parakeet TDT 0.6B v3

    70
  5. 5

    Inworld STT 1

    83
  6. 6

    Deepgram Nova 3

    99
  7. 7

    Deepgram Nova 2

    101
  8. 8

    Cartesia Ink 2

    108
  9. 9

    ElevenLabs Scribe v2 Realtime

    120
  10. 10

    AssemblyAI Universal 3.5 Pro

    146 ms
Figure 4. A5S V2 mean TTFS from the independent 947-item device trials inserted into Coval's rolling 30-day leaderboard as reported on August 29, 2026. A5S V2 uses Coval's metric definition and a released Coval evaluation set, but its one-off device trials are not official rolling leaderboard entries; the displayed positions are therefore hypothetical. Comparison cutoff: September 1, 2026.

Measurement protocol

TTFS follows Coval's definition: elapsed time from the reference end of speech to receipt of the final transcription segment. Each trial used 100 ms real-time audio pacing and concurrency one. Reference endpoint annotations isolate model finalization and exclude live endpoint detection, model loading, and warm-up.

DeviceItemsTTFS meanTTFS p50TTFS p95
MacBook Air, M494743.9945.5053.28
iPhone 16, A1894750.3950.0157.23

Table 4. TTFS in milliseconds, item-weighted across Coval stt-v1 and stt-v3. Lower is better.

MacBook Air · M4

947-item Coval trial

45.50 msmedian TTFS
Live execution on the recorded M4 MacBook Air. The terminal reports live TTFS using a configured 250 ms silence threshold.

iPhone 16 · A18

947-item Coval trial

50.01 msmedian TTFS
Live execution on the recorded iPhone 16. The application performs inference locally and reports live TTFS using the same silence threshold.

The live recordings use a 250 ms silence threshold to identify the end of speech. Their displayed values are implementation demonstrations and are distinct from the oracle-endpoint measurements in Table 4.

Pareto frontier: latency vs accuracy (Coval)

Lower-left is better

3%4%5%6%7%0200400600800 msWord error rateTime to Final Segment A5S V2 · Mac / iPhone44 / 50 ms · 4.11% AssemblyAI146 ms · 3.2% Scribe v2120 ms · 5.5% Nova-399 ms · 6.3% Google Chirp 3813 ms · 4.1% GPT-4o Transcribe742 ms · 4.7%
Figure 5. A5S V2 uses the complete 947-item Mac and iPhone trials, shown separately. Proprietary-cloud system points are Coval's 30-day averages as of August 29, 2026. The datasets and metric definitions are compatible, but the A5S V2 trials were run independently and are not official rolling leaderboard entries.

A locally executed speech model avoids an audio upload and continues to operate without a network connection. It also removes the provider's per-minute inference charge. These properties are separate from accuracy and are not represented in WER.

The proprietary cloud systems with public list prices publish rates ranging from $0.20 to $1.02 per audio hour under the plans shown below. A5S V2 has no external API charge per processed minute, although local compute, energy, integration, and hardware costs remain.

System used in evaluationDeployment in this comparisonPublished usage price
A5S V2Local, Mac or iPhoneNo external per-minute charge
xAI Grok Speech-to-Text StreamingProprietary cloud API$0.20 / audio hour
Deepgram Nova-3Proprietary cloud API$0.29 / audio hour
ElevenLabs Scribe v2 RealtimeProprietary cloud API$0.39 / audio hour
AssemblyAI Universal-3.5 ProProprietary cloud API$0.45 / audio hour
Google Gemini 3.5 Transcribe LiveProprietary cloud API~$0.54 / audio hour
Google Chirp 3Proprietary cloud API$0.96 / audio hour
OpenAI GPT Live TranscribeProprietary cloud API$1.02 / audio hour
Meta Muse TranscribeProprietary cloud APINot publicly listed

Table 5. Public list or pay-as-you-go prices checked in September 2026. Gemini's blended rate is the provider's token-based estimate. Discounts, add-ons, and enterprise terms vary.

Release v1 reports one saved inference trial for each system and corpus. Confidence intervals are not included. The planned analysis uses paired cluster bootstrap resampling by utterance for Mega-ASR and by complete meeting for the other corpora.

The Mega-ASR subset was sampled from a public training split because the standard test set did not separate robust systems in our preliminary experiments. A5S V2 was not trained, fine-tuned, selected, or prompted on the 1,250 chosen recordings or transcripts. We cannot determine whether third-party systems encountered these items or related upstream data. Mega-ASR should be read as a developer-created robustness diagnostic alongside the three meeting corpora, not as a speaker-disjoint test set.

These results measure English streaming transcription without diarization. They do not establish performance for other languages, speaker attribution, live endpoint detection, or downstream semantic tasks. Results are specific to the fixed data, normalization, provider versions, and request configurations in the release.