activity
20192026
most citedEnd-to-end spoofing detection with raw waveform CLDNNs

67 citations · 139 across the 28 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS2026

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

Xingwei Sun, Heinrich Dinkel, Gang Li +7

Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipel…

eess.AS2026

ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding

Yadong Niu, Tianzi Wang, Heinrich Dinkel +6

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in t…

eess.AS2025

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

Yadong Niu, Tianzi Wang, Heinrich Dinkel +7

While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because curren…

eess.AS2025

Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders

Xingwei Sun, Heinrich Dinkel, Yadong Niu +3

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal predictio…

eess.AS2020★ 67 cited

End-to-end spoofing detection with raw waveform CLDNNs

Heinrich Dinkel, Nanxin Chen, Yanmin Qian +1

Albeit recent progress in speaker verification generates powerful models, malicious attacks in the form of spoofed speech, are generally not coped with. Recent results in ASVSpoof2…