9 papers · 1 filter
Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
Shengfan Shen, Di Wu, Xingchen Song +5
Expressive speech synthesis for voice assistants requires flexible style control that adapts to explicit requests and broader interaction context. We propose Harness TTS, a lightwe…
Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a singl…
MiDashengLM: Efficient Audio Understanding with General Audio Captions
Heinrich Dinkel, Gang Li, Jizhong Liu +7
Current approaches for large audio language models (LALMs) often rely on closed data sources or proprietary models, limiting their generalization and accessibility. This paper intr…
Iterate to Differentiate: Enhancing Discriminability and Reliability in Zero-Shot TTS Evaluation
Shengfan Shen, Di Wu, Xingchen Song +5
Reliable evaluation of modern zero-shot text-to-speech (TTS) models remains challenging. Subjective tests are costly and hard to reproduce, while objective metrics often saturate,…
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
Heinrich Dinkel, Jiahao Zhou, Guanbo Wang +8
This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders…
GLAP: General contrastive audio-text pretraining across domains and languages
Heinrich Dinkel, Zhiyong Yan, Tianzi Wang +7
Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in Eng…