1 citations · 1 across the 2 of their papers we have counts for
36 papers
SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios
Jinming Zhang, Wei Rao, Xionghu Zhong +1
Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall performance. For instance, app…
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
Donghang Wu, Haoyang Zhang, Chen Chen +8
Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user…
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
Haoyang Li, Xuyi Zhuang, Azmat Adnan +6
Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech. We propose GenT…
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
Yukun Chen, Tianrui Wang, Zhaoxi Mu +2
High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealist…
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
Ziyang Ma, Ruiyang Xu, Zhenghao Xing +9
Fine-grained perception of multimodal information is critical for advancing human-AI interaction. With recent progress in audio-visual technologies, Omni Language Models (OLMs), ca…
Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
Nikita Kuzmin, Tao Zhong, Jiajun Deng +6
End-to-end full-duplex speech models feed user audio through an always-on LLM backbone, yet the speaker privacy implications of their hidden representations remain unexamined. Foll…