1 citations · 1 across the 7 of their papers we have counts for
7 papers · 1 filter
VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6
As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overloo…
The 10th AI City Challenge
Zheng Tang, Shuo Wang, David C. Anastasiu +34
The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with v…
From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning
Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat +5
Detecting a traffic anomaly does not establish whether a video-language model can explain what happened, localize it in time, or identify its causes. We introduce TAR (Traffic Anom…
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
Yatai Ji, An-Chieh Cheng, Yang Fu +13
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains…
Grounded 3D-Aware Spatial Vision-Language Modeling
An-Chieh Cheng, Yang Fu, Yatai Ji +12
We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding-…
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
Han Zhang, Wanting Jiang, Tomasz Kornuta +2
Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what…