activity
20232026
most citedFunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

9 citations · 19 across the 15 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Enabling Proactive Spoken Turns via a Generalized Style-Aware Full-Duplex Framework

Tianrui Pan, Qinglin Zhang, Chong Deng +6

Compared with half-duplex dialogue systems where the system waits for user turn completion before it responds, natural full-duplex dialogue systems require agents to act proactivel…

cs.CL2026

GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling

Hao-Xiang Xu, Chong Deng, Jiaqing Liu +5

Large Language Models (LLMs) extend their capabilities through function-calling (FC), which relies on training data with high quality, diversity, and broad coverage of scenario. Ho…

cs.CL20261 cited

Fun-Audio-Chat Technical Report

Tongyi Fun Team, Qian Chen, Luyao Cheng +10

Recent advancements in joint speech-text models show great potential for seamless voice interactions. However, existing models face critical challenges: temporal resolution mismatc…

cs.CL2025

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

Chao-Hong Tan, Qian Chen, Wen Wang +14

Recent studies on end-to-end (E2E) speech generation with large language models (LLMs) have attracted significant community attention, with multiple works extending text-based LLMs…

cs.CL20253 cited

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Qian Chen, Yafeng Chen, Yanni Chen +33

Recent advancements in large language models (LLMs) and multimodal speech-text models have laid the groundwork for seamless voice interactions, enabling real-time, natural, and hum…

cs.CL2024

OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Qinglin Zhang, Luyao Cheng, Chong Deng +8

Full-duplex spoken dialogue systems significantly surpass traditional turn-based dialogue systems, as they allow simultaneous bidirectional communication, closely mirroring human-h…