1 citations · 1 across the 4 of their papers we have counts for
9 papers
LaSR: Context-Aware Speech Recognition via Latent Reasoning
Heyang Liu, Ziyang Cheng, Jiayi Huang +5
Recent advances in Speech Large Language Models (Speech LLMs) have significantly enhanced spoken language understanding and reasoning. However, their contextual awareness is limite…
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
Ke Xu, Yuhao Wang, Ziyang Cheng +3
Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams…
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
Ziyang Cheng, Yuhao Wang, Heyang Liu +4
Recent Speech Large Language Models~(LLMs) have achieved impressive capabilities in end-to-end speech interaction. However, the prevailing autoregressive paradigm imposes strict se…
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
Heyang Liu, Yuhao Wang, Ziyang Cheng +7
Speech large language models (SpeechLLMs) have extended human-machine interactions from the text modality to the dynamic speech domain. Spoken dialogues convey diverse information,…
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
Heyang Liu, Ziyang Cheng, Yuhao Wang +6
The development of multi-modal large language models (LLMs) leads to intelligent approaches capable of speech interactions. As one of the most widely spoken languages globally, Man…
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
Yuhao Wang, Ziyang Cheng, Heyang Liu +4
Current end-to-end spoken language models (SLMs) have made notable progress, yet they still encounter considerable response latency. This delay primarily arises from the autoregres…