activity
20242026
most citedSRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation

3 citations · 5 across the 9 of their papers we have counts for

collaborators

11 papers

cs.CL2026

VLX-VR: An Agentic-Aware Video Reasoning Model

Sheng Li, Peng Liu, Qianqian Zhang +1

Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and singl…

cs.CL2026

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Kangjia Zhao, Jiajun Li, Haozhan Shen +8

Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass clos…

cs.AI2026

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Yutao Sun, Yanting Miao, Hao-Xuan Ma +8

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adap…

cs.CV2026

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

Tengjie Lin, Yutao Sun, Jingwei Ni +7

Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus o…

cs.CV2026

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schem…

cs.CV2026

MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning

Haozhan Shen, Shilin Yan, Hongwei Xue +5

Multimodal Large Language Models (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step depends on verified visual compositional c…