most citedOmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

1 citations · 1 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain +6

As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric consumer video. This overloo…

cs.CV2026

The 10th AI City Challenge

Zheng Tang, Shuo Wang, David C. Anastasiu +34

The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with v…

cs.CV2026

From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning

Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat +5

Detecting a traffic anomaly does not establish whether a video-language model can explain what happened, localize it in time, or identify its causes. We introduce TAR (Traffic Anom…

cs.CV2026

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models

Yatai Ji, An-Chieh Cheng, Yang Fu +13

Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains…

cs.CV2026

Grounded 3D-Aware Spatial Vision-Language Modeling

An-Chieh Cheng, Yang Fu, Yatai Ji +12

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding-…

cs.CV2026

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

Han Zhang, Wanting Jiang, Tomasz Kornuta +2

Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what…