8 papers
Direct Simultaneous Translation Activation for Large Audio-Language Models
Pei Zhang, Yiming Wang, Jialong Tang +4
Simultaneous speech-to-text translation (Simul-S2TT) aims to translate speech into target text in real time, outputting translations while receiving source speech input, rather tha…
BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models
Yujie Lin, Jiayao Ma, Qingguo Hu +5
Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interventions often adopt a difference-…
Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis
Han Wu, Yanming Sun, Yunhe Yang +1
Multimodal sentiment analysis (MSA) leverages information fusion from diverse modalities (e.g., text, audio, visual) to enhance sentiment prediction. However, simple fusion techniq…
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
Pei Zhang, Andong Chen, Xi Chen +3
Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is a…
Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation
Yiming Wang, Pei Zhang, Baosong Yang +2
LLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability. In this research track, we propose the C…
Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning
Yiming Wang, Pei Zhang, Baosong Yang +3
Real-world data deviating from the independent and identically distributed (i.i.d.) assumption of in-distribution training data poses security threats to deep networks, thus advanc…