3 papers
cs.AI2026
M-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering
Peijin Xie, Zhen Xu, Bingquan Liu +1
Multimodal large language models have recently shown promising progress in visual mathematical reasoning. However, their performance is often limited by a critical yet underexplore…
cs.CL2026
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
Shun Qian, Bingquan Liu, Chengjie Sun +2
The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-of-thought capabilities. However, there a…
cs.CV2024
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
Shun Qian, Bingquan Liu, Chengjie Sun +2
The projector plays a crucial role in multi-modal language models (MLLMs). The number of visual tokens it outputs affects the efficiency of the MLLM, while the quality of the visua…