8 papers
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching
Xingwei Sun, Heinrich Dinkel, Gang Li +7
Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipel…
SpeakerCard-1M: An Evidence-Grounded Corpus for In-the-Wild Speaker Verification
Junyi Peng, OldÅich Plchot, Xiao Song +9
Modern speaker verification (SV) systems rely on speaker embeddings that are effective but difficult to interpret or query in natural language. Most existing speech-text corpora ta…
Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation
Zheng Wang, Xiaobin Rong, Hang Su +6
Language model (LM)-based speech enhancement (SE) can generate natural-sounding speech, but under severe noise it often suffers from unreliable conditioning, leading to perceptuall…
Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding
Mingchen Shao, Hang Su, Wenjie Tian +6
While Large Audio Language Models (LALMs) achieve strong performance on short audio, they degrade on long-form inputs. This degradation is more severe in temporal awareness tasks,…
Borderless Long Speech Synthesis
Xingchen Song, Di Wu, Dinghao Zhou +12
Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both a…
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
Ziling Huang, Junnan Wu, Lichun Fan +4
Target speech extraction (TSE) has achieved strong performance in relatively simple conditions such as one-speaker-plus-noise and two-speaker mixtures, but its performance remains…