9 citations · 9 across the 2 of their papers we have counts for
3 papers
cs.SD2026
STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity
Sitong Cheng, Weizhen Bian, Songjun Cao +9
Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news reporting vs. dramatic dialogue),…
cs.SD2025
Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
Hualei Wang, Yiming Li, Shuo Ma +2
Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately…
eess.AS2024★ 9 cited
Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
Yiming Li, Zhifang Guo, Xiangdong Wang +1
Recent advances have been witnessed in audio-language joint learning, such as CLAP, that shows much success in multi-modal understanding tasks. These models usually aggregate uni-m…