10 papers · 1 filter
GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark
Yujie Tu, Yifan Yang, Tianrui Wang +36
While modern ASR systems achieve low error rates on high-resource benchmarks, such performance often overestimates real-world robustness. Existing evaluations address challenges in…
EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning
Siyuan Zhang, Jian Zong, Junyu Wang +7
While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with c…
Audio Imitator: Controlling Timbre and Tempo in Video2Audio Synthesis with Audio Reference
Jiahui Zhao, Tianrui Wang, Chunyu Qiang +4
Video-to-audio generation has made significant progress in achieving semantic consistency and temporal alignment from silent videos. However, audio contains rich stylistic attribut…
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang, Ziyang Ma, Yizhou Peng +26
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its con…
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
Chunyu Qiang, Xiaopeng Wang, Kang Yin +11
Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous…
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
Tianrui Wang, Meng Ge, Cheng Gong +11
While LLM-based TTS models exhibit zero-shot emotion and speaker cloning, their cloning fidelity and pronunciation clarity degrade on unseen domains. Fine-tuning is essential for a…