activity
20242026
collaborators

11 papers

cs.AI2026

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Yutao Sun, Yanting Miao, Hao-Xuan Ma +8

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adap…

cs.CV2026

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

Tengjie Lin, Yutao Sun, Jingwei Ni +7

Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus o…

cs.CV2026

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schem…

cs.CV2026

Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning

Zilun Zhang, Zian Guan, Tiancheng Zhao +7

Referring expression understanding in remote sensing poses unique challenges, as it requires reasoning over complex object-context relationships. While supervised fine-tuning (SFT)…

cs.CV2026

DetailCLIP: Injecting Image Details into CLIP's Feature Space

Zilun Zhang, Cuifeng Shen, Yuan Shen +4

Although CLIP-like Visual Language Models provide a functional joint feature space for image and text, due to the limitation of the CILP-like model's image input size (e.g., 224),…

cs.CV2026

MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning

Haozhan Shen, Shilin Yan, Hongwei Xue +5

Multimodal Large Language Models (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step depends on verified visual compositional c…