10 papers
Culture In a Frame: CB as a Comic-Based Benchmark for Multimodal Culturally Awareness
Yuchen Song, Andong Chen, Wenxin Zhu +4
Cultural awareness capabilities have emerged as a critical capability for Multimodal Large Language Models (MLLMs). However, current benchmarks lack progressed difficulty in their…
Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling
Andong Chen, Wenxin Zhu, Qiuyu Ding +3
Chain-of-Thought reasoning has driven large language models to extend from thinking with text to thinking with images and videos. However, different modalities still have clear lim…
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
Wenxin Zhu, Andong Chen, Yuchen Song +4
With the remarkable success of Multimodal Large Language Models (MLLMs) in perception tasks, enhancing their complex reasoning capabilities has emerged as a critical research focus…
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
Pei Zhang, Andong Chen, Xi Chen +3
Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is a…
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
Sirui Wang, Andong Chen, Tiejun Zhao
Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predef…
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
Henglyu Liu, Andong Chen, Kehai Chen +4
Recent advancement of large language models (LLMs) has led to significant breakthroughs across various tasks, laying the foundation for the development of LLM-based speech translat…