1 citations · 1 across the 7 of their papers we have counts for
11 papers
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
Qingjian Lin, Yuxin Li, Haoyang Zhang +14
Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale languag…
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
Haoyang Zhang, Jun Chen, Donghang Wu +13
Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses.…
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
Vision Foundation Models as Generalist Tokenizers for Image Generation
Anlin Zheng, Qi Han, Xin Wen +5
In this work, we explore the largely unexplored direction of building a generalist image tokenizer directly on top of a frozen vision foundation model (VFM). To build this tokenize…
Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
Houyi Li, Ka Man Lo, Shijie Xuyang +7
Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense…
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
Che Liu, Lichao Ma, Xiangyu Tony Zhang +4
Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer…