1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.SD2025★ 1 cited
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
Yixuan Hou, Heyang Liu, Yuhao Wang +5
Thanks to the steady progress of large language models (LLMs), speech encoding algorithms and vocoder structure, recent advancements have enabled generating speech response directl…
cs.CV2025
DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor
Yan Zhao, Zhengxue Cheng, Junxuan Zhang +3
Most learning-based lossless compressors are designed for a single modality, requiring separate models for multi-modal data and lacking flexibility. However, different modalities v…
cs.CL2025★ 1 cited
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
Yuhao Wang, Heyang Liu, Ziyang Cheng +4
Speech large language models (LLMs) have emerged as a prominent research focus in speech processing. We introduce VocalNet-1B and VocalNet-8B, a series of high-performance, low-lat…