8 papers · 1 filter
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
Wenxin Zhu, Andong Chen, Yuchen Song +4
With the remarkable success of Multimodal Large Language Models (MLLMs) in perception tasks, enhancing their complex reasoning capabilities has emerged as a critical research focus…
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
Pei Zhang, Andong Chen, Xi Chen +3
Large language models (LLMs) have expanded from text to speech, giving rise to Speech Large Models (SLMs) that support recognition, translation, and synthesis. A key challenge is a…
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
Sirui Wang, Andong Chen, Tiejun Zhao
Emotional text-to-speech (E-TTS) is central to creating natural and trustworthy human-computer interaction. Existing systems typically rely on sentence-level control through predef…
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
Henglyu Liu, Andong Chen, Kehai Chen +4
Recent advancement of large language models (LLMs) has led to significant breakthroughs across various tasks, laying the foundation for the development of LLM-based speech translat…
Evaluating o1-Like LLMs: Unlocking Reasoning for Translation through Comprehensive Analysis
Andong Chen, Yuchen Song, Wenxin Zhu +4
The o1-Like LLMs are transforming AI by simulating human cognitive processes, but their performance in multilingual machine translation (MMT) remains underexplored. This study exam…
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation
Andong Chen, Yuchen Song, Kehai Chen +3
Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sente…