SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing
arXiv:2609.26255
Abstract
Text-to-speech systems are increasingly used for drama dubbing, yet evaluation protocols remain sentence-level, leaving critical scene-level behaviors insufficiently measured. We present SceneTTS-Bench, a benchmark that evaluates TTS along three dimensions: timbre consistency across character turns, emotional expressiveness on high-tension utterances, and rhythm coherence under segmented long-form synthesis. The corpus comprises real-world and generated drama scripts totaling 160 bilingual scenes and approximately 10,300 utterances, with real-world scripts serving as the primary source (100 scenes) and generated scripts as a supplementary source (60 scenes), demonstrating the framework's extensibility through synthetic data augmentation. A backend-agnostic Canonical Intermediate Representation ensures fair cross-system comparison. Three automatic pipelines produce per-utterance diagnostics: Speaker Consistency Score for timbre-drift detection, Under-Acting Ratio for under-acting identification, and Rate Discontinuity Ratio for rate-discontinuity quantification. Experiments on four TTS systems confirm that each system exhibits distinct weaknesses and that scene-level rankings diverge substantially from sentence-level metrics. Benchmark resources are publicly available at https://piedpiperg.github.io/scenetts-bench/ .
7 pages, 3 figures; submitted to the ACM Multimedia (ACM MM) 2026 Dataset Track