12 papers · 1 filter
Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text
Jiahao Mei, Heinrich Dinkel, Yadong Niu +7
Audio generation has long been fragmented, with speech, music, and sound effects produced by domain-specific models that fail to jointly generate coherent audio scenes from a singl…
DashengTokenizer: One layer is enough for unified audio understanding and generation
Heinrich Dinkel, Xingwei Sun, Gang Li +8
This paper introduces DashengTokenizer, a continuous audio tokenizer engineered for joint use in both understanding and generation tasks. Unlike conventional approaches, which trai…
MiDashengLM: Efficient Audio Understanding with General Audio Captions
Heinrich Dinkel, Gang Li, Jizhong Liu +7
Current approaches for large audio language models (LALMs) often rely on closed data sources or proprietary models, limiting their generalization and accessibility. This paper intr…
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
Heinrich Dinkel, Jiahao Zhou, Guanbo Wang +8
This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders…
GLAP: General contrastive audio-text pretraining across domains and languages
Heinrich Dinkel, Zhiyong Yan, Tianzi Wang +7
Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in Eng…
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
Junbo Zhang, Heinrich Dinkel, Yadong Niu +4
We introduces X-ARES (eXtensive Audio Representation and Evaluation Suite), a novel open-source benchmark designed to systematically assess audio encoder performance across diverse…