2 papers
cs.CL2025
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models
Bo Li, Chengben Xu, Wufeng Zhang
This paper presents Seewo's systems for both tracks of the Multilingual Conversational Speech Language Model Challenge (MLC-SLM), addressing automatic speech recognition (ASR) and…
cs.CV2025
Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
Junrong Yue, Yifan Zhang, Chuan Qin +5
Vision-and-Language Navigation (VLN) aims to enable embodied agents to follow natural language instructions and reach target locations in real-world environments. While prior metho…