activity
20242026
collaborators

8 papers

cs.SD2026

Direct Simultaneous Translation Activation for Large Audio-Language Models

Pei Zhang, Yiming Wang, Jialong Tang +4

Simultaneous speech-to-text translation (Simul-S2TT) aims to translate speech into target text in real time, outputting translations while receiving source speech input, rather tha…

cs.AI2025

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

Yujie Lin, Jiayao Ma, Qingguo Hu +5

Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interventions often adopt a difference-…

cs.LG2025

Beyond Simple Fusion: Adaptive Gated Fusion for Robust Multimodal Sentiment Analysis

Han Wu, Yanming Sun, Yunhe Yang +1

Multimodal sentiment analysis (MSA) leverages information fusion from diverse modalities (e.g., text, audio, visual) to enhance sentiment prediction. However, simple fusion techniq…

cs.CL2025

PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs

Pei Zhang, Andong Chen, Xi Chen +3

Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is a…

cs.CL2025

Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation

Yiming Wang, Pei Zhang, Baosong Yang +2

LLM self-evaluation relies on the LLM's own ability to estimate response correctness, which can greatly improve its deployment reliability. In this research track, we propose the C…

cs.CL2024

Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning

Yiming Wang, Pei Zhang, Baosong Yang +3

Real-world data deviating from the independent and identically distributed (i.i.d.) assumption of in-distribution training data poses security threats to deep networks, thus advanc…