5 citations · 24 across the 44 of their papers we have counts for
26 papers · 1 filter
TRAM: Enhancing Multimodal Reasoning with Trajectory-Derived Auxiliary Memory
Kang Liu, Zijing Wang, Yongkang Liu +5
Multimodal Large Reasoning Models (MLRMs) have achieved strong performance on tasks requiring visual understanding and multi-step inference. However, as reasoning trajectories grow…
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue
Wen Zhang, Xiaocui Yang, Zhuoyue Gao +3
Empathetic spoken dialogue systems require not only semantically appropriate responses but also emotionally aligned prosodic expression. However, cascade pipelines often discard ac…
GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification
Faxian Wan, Xiaocui Yang, Yifan Cao +3
Moving beyond the traditional binary classification paradigm of Multimodal Sarcasm Detection, Multimodal Sarcasm Target Identification (MSTI) presents a more formidable challenge,…
NEAT: Neuron-Based Early Exit for Large Reasoning Models
Kang Liu, Yongkang Liu, Xiaocui Yang +5
Large Reasoning Models (LRMs) often suffer from \emph{overthinking}, a phenomenon in which redundant reasoning steps are generated after a correct solution has already been reached…
PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs
Zijing Wang, Yongkang Liu, Mingyang Wang +8
Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically de…
CIRAG: Construction-Integration Retrieval and Adaptive Generation for Multi-hop Question Answering
Zili Wei, Xiaocui Yang, Yilin Wang +5
Triple-based Iterative Retrieval-Augmented Generation (iRAG) mitigates document-level noise for multi-hop question answering. However, existing methods still face limitations: (i)…