3 citations · 5 across the 9 of their papers we have counts for
11 papers
VLX-VR: An Agentic-Aware Video Reasoning Model
Sheng Li, Peng Liu, Qianqian Zhang +1
Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and singl…
Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling
Kangjia Zhao, Jiajun Li, Haozhan Shen +8
Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass clos…
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
Yutao Sun, Yanting Miao, Hao-Xuan Ma +8
Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adap…
REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation
Tengjie Lin, Yutao Sun, Jingwei Ni +7
Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than genuine visual ability, thus o…
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1
Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schem…
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
Haozhan Shen, Shilin Yan, Hongwei Xue +5
Multimodal Large Language Models (MLLMs) are increasingly used to carry out visual workflows such as navigating GUIs, where the next step depends on verified visual compositional c…