paper

VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition

arXiv:2608.28916

Abstract

Automatic speech recognition is usually evaluated with word error rate (WER), although voice workflows often require exact written values. VoiceCodeBench measures whether transcripts preserve identifiers, paths, commands, and other structured tokens needed by downstream software. It contains 300 human-recorded English workplace segments (5.59 hours, 85 speakers) and 1,482 audited entities across 26 types and eight domains. Under a raw-audio-only protocol, we evaluate 19 batch and streaming systems using WER, Canonical Token/Entity Match (CTEM), and strict segment-level Task Success Rate (TSR). Across systems, WER has little rank agreement with CTEM (Spearman ) or TSR (). The best CTEM and TSR are 91.8% and 68.7%. Even the strongest system therefore leaves nearly one-third of recordings with an unrecovered critical value. Symbol-, separator-, and boundary-sensitive entities account for most errors.

7 pages, 1 figure, 9 tables

VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition · wovepaper