28 papers
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Ming Wang, Yuqing Zhang, Tingna Xie +5
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals…
TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory
Kang Liu, Zijing Wang, Yongkang Liu +5
Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as reasoning trajectories grow…
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue
Wen Zhang, Xiaocui Yang, Zhuoyue Gao +3
Empathetic spoken dialogue systems require not only semantically appropriate responses but also emotionally aligned prosodic expression. However, cascade pipelines often discard ac…
Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models
Yichen Gao, Yiqun Zhang, Zijing Wang +7
Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavail…
SAILRec: Steering LLM Attention to Dual-Side Semantically Aligned Collaborative Embeddings for Recommendation
Xi Wu, Jiale Wang, Zihan Wang +5
Recent LLM-based recommenders enhance language models with collaborative embeddings from user-item interactions, but making such embeddings available does not ensure their proper u…
CR-Seg: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation
Yifan Cao, Xiaocui Yang, Faxian Wan +3
Reasoning segmentation aims to segment target objects described by complex language through joint visual-textual reasoning. Existing methods typically rely on either learned semant…