2 papers
cs.AI2026
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
Yutao Sun, Yanting Miao, Hao-Xuan Ma +8
Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adap…
cs.CV2026
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
Yanting Miao, Yutao Sun, Dexin Wang +8
Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However…