From the 1 of 9 linked papers with an AI index.
9 papers
Beyond the Eye: Efficient Multimodal Reasoning via Self-Regulated Implicit Visual Tools
Xiuwei Chen, Quanlin Chen, Wentao Hu +8
The paper introduces Beyond the Eye (BEE), an implicit visual‑tool framework for multimodal large language models that learns to self‑regulate when to invoke visual tools, reducing…
SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning
Xiuwei Chen, Wentao Hu, Hanhui Li +9
Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing MLLMs necessitates high-quality vision…
Latent Visual States for Efficient Multimodal Reasoning
Xiuwei Chen, Wentao Hu, Yongxin Wang +8
The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete out…
Intend, Reflect, Refine: An Adaptive Multimodal Reflection Framework for Autonomous Driving
Zisheng Chen, Yuping Qiu, Jianhua Han +6
Recent Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving by incorporating reasoning for better interpretability and planning quality. However, most ex…
AtomicVLA: Unlocking the Potential of Atomic Skill Learning in Robots
Likui Zhang, Tao Tang, Zhihao Zhan +9
Recent advances in Visual-Language-Action (VLA) models have shown promising potential for robotic manipulation tasks. However, real-world robotic tasks often involve long-horizon,…
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
Zisheng Chen, Chunwei Wang, Runhui Huang +6
In this paper, we introduce SemHiTok, a unified image Tokenizer via Semantic-Guided Hierarchical codebook that provides consistent discrete representations for multimodal understan…