Industry Partner

VoiceCodeBench

Updated 8/27/2026

How good are speech-to-text models at handling exact structured values in English workplace speech?

Partners in Evaluation

Key Takeaways

  • Speech-to-text (STT) still fails on values that contain separators and punctuation: URLs (46% mean recovery), commands (51%), file paths (52%), postal addresses (57%), email addresses (62%).
  • STT can effectively handle dates, plain numbers, percentages, measurements and phone extensions, all 98-99%.
  • Most accurate model is GPT Live Transcribe, 67.67% Task Success Rate (TSR) and 91.36% Canonical Token/Entity Match (CTEM).
  • Best realtime pick: GPT Live Transcribe has a fast time-to-first-partial latency (2.52s); Deepgram Nova-3 (Streaming) is the fastest accurate option (2.17s, 59.33%).
  • Cheapest offline option: Whisper Large-v3 Turbo, 57.00% TSR at $0.0007 per task, roughly a twenty-fifth of the leader’s cost.

Background

VoiceCodeBench is a benchmark designed by besimple.ai that asks which speech-to-text system is best at identifying the key terms that critical workflows such as phone agents and dictation apps depend on. STT is used by voice AI platforms to fill API arguments: the address an order ships to, the email a summary goes to, the command an agent runs. A system that does well identifying low-risk vocab can still perform poorly in real-world workflows, as the transcript may look correct but contain a structural error, so the failure reaches the downstream system as a wrong tool argument or an invalid database field rather than as an obvious transcription error.

Capturing these terms is what proves high-stakes reliability in production. What makes VoiceCodeBench unique is that its scores represent not just Word Error Rate (WER) but also Task Success Rate (TSR), which a model only passes if it gets every target entity inside of a recording correct. There is also Canonical Token/Entity Match (CTEM), which is the share of target entities that can be recovered from the Automatic Speech Recognition (ASR) transcript. CTEM matters because you may be able to programmatically salvage an incorrect transcription.

The dataset is 300 human-recorded English workplace clips, 5.6 hours of audio between 34.9 and 122.9 seconds long, holding 1,482 audited target entities across 26 entity types and 8 workflow domains. It was built entity-first: the entity types and difficulty band were chosen per recording before the transcript was written around them, so coverage is controlled rather than whatever structured values happen to appear in free-form speech. Every structured value is synthetic, and the benchmark is test-only with no training corpus.


Results

The main metric is TSR, which is the share of recordings where the STT model correctly identified all of the target entities. Recoverable renderings of each canonical answer are also marked correct: a phone number read out digit by digit and written back as an unbroken digit string still counts, while a lost separator, a dropped underscore or a substituted unit does not. Unlike CTEM, TSR does not provide partial credit. This means that if 2 of 3 target entities are correctly identified, CTEM would be 66% while TSR would be 0%. Recordings carry 4.94 target entities on average and up to 8, so a single corrupted value fails the recording. CTEM is calculated across all 1,482 target entities inside of the benchmark.

Below is an example of what is passed into the ASR model, and what it can produce.

Public transcription examples

Compare how GPT-4o recovers exact entities across low, mid, and high performance samples from the public benchmark.

retail logistics order · return authorization

retail_logistics_order_032 · 58.14s

Low performance

1 of 4 entities recovered · 25%

url: Missedacronym or initialism: Missedproduct code: Missedspelled sequence: Recovered

Expected

Canonical reference

The inspection photos are attached, but the return reason needs a quick correction before the warehouse accepts the carton. Use https://returns.lakeroute.example.org/rma/LR-2105/inspection and keep the authorization under RMA-WRONG. The serial on the unit is SN-K8R-2001, and the suffix on the taped tag is spelled U W Y 1 2 0. Sorry, the issue is that the first reviewer called it a damaged return, but the customer received the wrong part in a sealed box. Please keep the existing authorization, move the photos into the inspection tab, and hold the pickup label until that suffix matches. Once it matches, release the label without creating a second credit review.

Observed

GPT-4o transcription output

The inspection photos are attached, but the return reason needs a quick correction before the warehouse accepts the carton. Use https://returns.lakeroot.example.org/rma/lr-2105/inspection and keep the authorization under RMA-WROMG. The serial on the unit is SN-KHR-2001, and the suffix on the DEP tag is spelled UWY120. Sorry, the issue is that the first reviewer called it a damaged return, but the customer received the wrong part in a sealed box. Please, keep the existing authorization, move the photos into the inspection tab, and hold the pickup label until that suffix matches. Once it matches, release the label without creating a second credit review.

Aggregate metric comparison

Side by side comparison of all models and metrics. WAcc (Word Accuracy) = 1 - WER.

TSRCTEMWAcc
GPT Live Transcribe
TSR
67.7%
CTEM
91.4%
WAcc
94.1%
Cartesia Ink 2
TSR
62.0%
CTEM
89.5%
WAcc
90.8%
GPT 4o Transcribe
TSR
59.7%
CTEM
88.7%
WAcc
96.0%
Deepgram Nova-3 (Streaming)
TSR
59.3%
CTEM
88.0%
WAcc
90.5%
Inworld STT 1
TSR
57.3%
CTEM
88.0%
WAcc
91.9%
Whisper Large-v3 Turbo
TSR
57.0%
CTEM
87.6%
WAcc
93.3%
Deepgram Flux General English
TSR
55.7%
CTEM
86.6%
WAcc
90.6%
Google Cloud Chirp 3
TSR
55.3%
CTEM
86.8%
WAcc
93.4%
ElevenLabs Scribe v2 Realtime
TSR
55.3%
CTEM
86.8%
WAcc
94.9%
Whisper Large-v3
TSR
54.3%
CTEM
87.2%
WAcc
94.2%
Voxtral Mini Transcribe 2
TSR
53.0%
CTEM
87.0%
WAcc
94.5%
Grok STT
TSR
52.0%
CTEM
84.7%
WAcc
91.0%
GPT 4o Mini Transcribe (Streaming)
TSR
51.7%
CTEM
86.6%
WAcc
95.2%
Voxtral Mini Transcribe Realtime 2602
TSR
46.0%
CTEM
83.3%
WAcc
93.4%
Cohere Transcribe
TSR
44.7%
CTEM
83.9%
WAcc
93.0%
Azure Speech Universal Language Model
TSR
37.3%
CTEM
80.7%
WAcc
89.6%
AssemblyAI Universal-3.5 Pro
TSR
30.3%
CTEM
77.1%
WAcc
94.5%

There are two types of ASR models, realtime and non-realtime. Realtime models stream the transcription as the audio arrives instead of waiting until all of the audio has been processed. Streaming enables realtime systems that need low latency to perform in production environments such as voice AI. For those systems the number that matters is time to first partial (TTFP), not total request time.

Average TTFP

Time from request/audio start to the first partial transcript. Models without a measured TTFP are omitted.

Voxtral Mini Transcribe Realtime 2602
0.66s
Azure Speech Universal Language Model
1.10s
Deepgram Flux General English
1.54s
GPT 4o Mini Transcribe (Streaming)
1.57s
AssemblyAI Universal-3.5 Pro
1.79s
Deepgram Nova-3 (Streaming)
2.17s
Grok STT
2.27s
Inworld STT 1
2.47s
GPT Live Transcribe
2.52s
ElevenLabs Scribe v2 Realtime
2.69s
Google Cloud Chirp 3
7.55s

Current weaknesses within ASR technology are in commands, email addresses, file paths, postal addresses, URLs, and environment variables, where separators, casing and punctuation are part of the value itself. Dates, times, plain numbers, percentages and measurements are effectively solved.

Entity recovery by type

Heatmap of aggregate exact-token recovery rate.

Modelaccount or record numberacronym or initialismcli flagcode symbolcommandcurrency amountdatedomain termemail addressenvironment variablefile pathip addressmeasurementpercentageperson or team namephone extensionphone numberplain numberport numberpostal addressproduct codereference idspelled sequencetimeurlversion
GPT Live Transcribe92.3%95.3%95.5%94.3%76.0%100.0%100.0%95.0%78.5%88.6%66.7%96.0%100.0%100.0%94.4%100.0%100.0%100.0%96.7%55.0%93.3%94.0%94.8%98.3%62.9%93.3%
Cartesia Ink 293.8%98.8%88.6%82.9%56.0%94.7%100.0%95.0%78.5%88.6%52.9%96.0%100.0%100.0%94.4%100.0%98.3%98.5%96.7%55.0%93.3%94.0%97.9%96.7%56.5%96.7%
GPT 4o Transcribe92.3%97.6%95.5%85.7%46.0%100.0%100.0%90.0%66.2%77.1%64.7%84.0%100.0%100.0%100.0%100.0%96.7%100.0%100.0%60.0%95.6%92.7%94.8%95.0%53.2%93.3%
Deepgram Nova-3 (Streaming)90.8%98.8%86.4%80.0%54.0%98.7%98.9%85.0%69.2%77.1%51.0%88.0%98.4%100.0%100.0%96.7%98.3%100.0%86.7%67.5%88.9%94.7%94.8%98.3%54.8%93.3%
Inworld STT 196.9%97.6%93.2%91.4%54.0%96.0%98.9%95.0%58.5%88.6%62.7%84.0%100.0%100.0%100.0%100.0%93.3%100.0%93.3%55.0%86.7%92.0%96.9%95.0%51.6%93.3%
Whisper Large-v3 Turbo92.3%100.0%86.4%82.9%40.0%98.7%98.9%90.0%69.2%85.7%49.0%92.0%100.0%100.0%94.4%100.0%95.0%96.9%86.7%55.0%92.2%92.0%96.9%96.7%56.5%96.7%
Deepgram Flux General English90.8%96.5%88.6%77.1%52.0%98.7%96.6%85.0%63.1%82.9%49.0%96.0%93.4%98.0%94.4%100.0%100.0%100.0%96.7%52.5%87.8%91.3%95.9%95.0%53.2%90.0%
Google Cloud Chirp 393.8%95.3%86.4%91.4%56.0%86.7%100.0%95.0%61.5%80.0%58.8%96.0%98.4%98.0%94.4%100.0%96.7%98.5%90.0%52.5%91.1%88.7%92.8%96.7%58.1%90.0%
ElevenLabs Scribe v2 Realtime90.8%98.8%84.1%88.6%44.0%92.0%98.9%100.0%70.8%51.4%43.1%96.0%100.0%100.0%100.0%96.7%95.0%100.0%96.7%60.0%87.8%96.0%96.9%98.3%46.8%96.7%
Whisper Large-v387.7%97.6%86.4%85.7%52.0%97.3%96.6%85.0%75.4%88.6%49.0%92.0%100.0%100.0%94.4%100.0%95.0%96.9%96.7%60.0%92.2%92.7%97.9%93.3%35.5%96.7%
Voxtral Mini Transcribe 290.8%98.8%90.9%85.7%58.0%97.3%98.9%90.0%66.2%88.6%58.8%88.0%100.0%96.0%94.4%96.7%98.3%98.5%93.3%55.0%94.4%90.7%96.9%95.0%22.6%93.3%
Grok STT93.8%96.5%90.9%88.6%54.0%92.0%97.8%95.0%63.1%74.3%47.1%48.0%100.0%96.0%94.4%100.0%95.0%98.5%93.3%50.0%83.3%89.3%92.8%93.3%45.2%93.3%
GPT 4o Mini Transcribe (Streaming)84.6%94.1%93.2%91.4%58.0%98.7%100.0%95.0%66.2%77.1%64.7%76.0%100.0%100.0%100.0%100.0%96.7%96.9%100.0%50.0%84.4%90.7%82.5%95.0%58.1%93.3%
Voxtral Mini Transcribe Realtime 260286.2%97.6%88.6%80.0%46.0%96.0%97.8%90.0%40.0%85.7%58.8%96.0%98.4%94.0%100.0%93.3%96.7%98.5%93.3%55.0%87.8%84.7%95.9%95.0%21.0%83.3%
Cohere Transcribe92.3%96.5%90.9%88.6%40.0%93.3%98.9%90.0%35.4%91.4%35.3%96.0%98.4%100.0%94.4%100.0%90.0%96.9%90.0%65.0%82.2%90.0%93.8%96.7%37.1%96.7%
Azure Speech Universal Language Model83.1%94.1%84.1%71.4%34.0%93.3%98.9%100.0%43.1%74.3%25.5%88.0%96.7%100.0%88.9%100.0%95.0%95.4%93.3%67.5%74.4%82.7%93.8%95.0%40.3%76.7%
AssemblyAI Universal-3.5 Pro72.3%96.5%86.4%85.7%44.0%85.3%95.5%95.0%44.6%68.6%47.1%72.0%98.4%98.0%100.0%96.7%76.7%95.4%86.7%47.5%74.4%67.3%80.4%98.3%32.3%90.0%

Methodology

We followed the same methodology reported inside of the VoiceCodeBench paper to ensure alignment. The code from Huggingface was ported over to Valkyrie, our own agent orchestrator. This change was only infrastructure and did not affect the scores. Models receive the raw audio only, with no benchmark prompting, entity hints, custom vocabulary, grammar constraints or post-ASR correction. Entity scoring uses the released LLM-assisted recoverability verifier, which the paper audited against a human reviewer on a stratified sample of 200 decisions with full agreement.

The metrics were also calculated the same way as described in the paper. We used the provided WER to calculate Word Accuracy (WAcc), 1 - WER, which shows how many words the model got correct. This is because it is easier to understand that models with a higher score performed better than models with a lower score. WER shows the opposite of that and can be misleading. Our WER is the format-invariant WER the released scorer computes, which accepts either the spoken or the written rendering of a target value, so it is not comparable to the stricter WER column published in the paper. Our own scores are also not directly comparable to the paper’s twelve baselines, since the model versions, providers and streaming settings may differ.

TTFP is defined as time to first partial and represents the time it took for the ASR model to produce a partial text chunk collected using realtime. Not all models support this and are omitted from visuals that compare TTFP. Latency is defined as the time it takes to transcribe the entire audio clip, which for realtime models includes streaming the audio in at speaking pace.

Cost is not tracked for models that do not publish a clear USD price per unit of audio, such as those billed in plan-dependent credits, in committed hourly capacity, or on dedicated instances.

The benchmark is English-only and its structured values are synthetic, so it measures exact-value recovery rather than multilingual or conversational transcription quality.

Citation (BibTeX)

@misc{voicecodebench2026,
title  = {VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition},
author = {Baumgartner, Tyler and Tai, Brandon and Kaelin-Martin, Lisa and Fan, Candice and Debaupte, Luc and Wang, Bill and Zhong, Yi},
year   = {2026},
note   = {Benchmark dataset and paper},
url    = {https://huggingface.co/datasets/besimple-ai/voice-code-bench}
}