From the 1 of 12 linked papers with an AI index.
12 papers
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools
Xiuwei Chen, Quanlin Chen, Wentao Hu +8
The paper introduces Beyond the Eye (BEE), an implicit visual‑tool framework for multimodal large language models that learns to self‑regulate when to invoke visual tools, reducing…
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
Latent Visual States for Efficient Multimodal Reasoning
Xiuwei Chen, Wentao Hu, Yongxin Wang +8
The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete out…
Regret Pre-training: Bridging Prior and Posterior Views for Enhanced Knowledge Grounding
Mingkuan Zhao, Xiayu Sun, Wentao Hu +5
Causal language models factorize sequence probabilities using only preceding context, leaving future information unexploited during training despite its availability in the trainin…
Hallucinations as Orthogonal Noise: Inference-Time Manifold Alignment via Dynamic Contextual Orthogonalization
Mingkuan Zhao, Wentao Hu, Tianchen Huang +6
Hallucination in Large Language Models (LLMs), characterized by the generation of content inconsistent with contextual facts or logical constraints -- remains a persistent challeng…
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
Mingkuan Zhao, Yide Gao, Wentao Hu +6
Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent fa…