9 citations · 9 across the 2 of their papers we have counts for
Showing eess.ASShow all
2 papers · 1 filter
eess.AS2026
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation
Zhen Ye, Xu Tan, Yiming Li +10
Spoken dialogue models typically start from text LLM backbones, yet reasoning often degrades when conditioning on speech instead of text. We attribute part of this modality gap to…
eess.AS2024★ 9 cited
Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
Yiming Li, Zhifang Guo, Xiangdong Wang +1
Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-m…