Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
ZhengXian Wu, Hangrui Xu, Kai Shi +8
Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fixed retrieve-then-generate pip…
cs.CV2026
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
Zhengxian Wu, Kai Shi, Chuanrui Zhang +10
Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-…
cs.CV2025
DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
Kai Shi, Jun Yang, Ni Yang +6
Mobile Phone Agents (MPAs) have emerged as a promising research direction due to their broad applicability across diverse scenarios. While Multimodal Large Language Models (MLLMs)…