3 papers
cs.AI2026
M-ACE: Rectifying Visual Perception in Multimodal Math Reasoning via Multi-Agentic Context Engineering
Peijin Xie, Zhen Xu, Bingquan Liu +1
Multimodal large language models have recently shown promising progress in visual mathematical reasoning. However, their performance is often limited by a critical yet underexplore…
cs.CV2025
TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
Peijin Xie, Shun Qian, Bingquan Liu +3
Document images encapsulate a wealth of knowledge, while the portability of spoken queries enables broader and flexible application scenarios. Yet, no prior work has explored knowl…
cs.CV2024
Expand VSR Benchmark for VLLM to Expertize in Spatial Rules
Peijin Xie, Lin Sun, Bingquan Liu +4
Distinguishing spatial relations is a basic part of human cognition which requires fine-grained perception on cross-instance. Although benchmarks like MME, MMBench and SEED compreh…