activity
20192025
most citedEnd-to-end spoofing detection with raw waveform CLDNNs

67 citations · 137 across the 11 of their papers we have counts for

collaborators

19 papers

eess.AS2025

Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders

Xingwei Sun, Heinrich Dinkel, Yadong Niu +3

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal predictio…

cs.SD2025

GLAP: General contrastive audio-text pretraining across domains and languages

Heinrich Dinkel, Zhiyong Yan, Tianzi Wang +7

Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains. Current CLAP methods enable sound and music retrieval in Eng…

cs.SD2025

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance

Junbo Zhang, Heinrich Dinkel, Yadong Niu +4

We introduces X-ARES (eXtensive Audio Representation and Evaluation Suite), a novel open-source benchmark designed to systematically assess audio encoder performance across diverse…

cs.SD2025

Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

Gang Li, Jizhong Liu, Heinrich Dinkel +3

Recently, reinforcement learning (RL) has been shown to greatly enhance the reasoning capabilities of large language models (LLMs), and RL-based approaches have been progressively…

cs.SD20251 cited

The ICME 2025 Audio Encoder Capability Challenge

Junbo Zhang, Heinrich Dinkel, Qiong Song +8

This challenge aims to evaluate the capabilities of audio encoders, especially in the context of multi-task learning and real-world applications. Participants are invited to submit…

cs.SD2022

An empirical study of weakly supervised audio tagging embeddings for general audio representations

Heinrich Dinkel, Zhiyong Yan, Yongqing Wang +2

We study the usability of pre-trained weakly supervised audio tagging (AT) models as feature extractors for general audio representations. We mainly analyze the feasibility of tran…