collaborators

9 papers

cs.CV2026

One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory

Chenhao Zheng, Jieyu Zhang, Mohammadreza Salehi +5

Effective video tokenization is critical for scaling transformer models for long videos. Current approaches tokenize videos using space-time patches, leading to excessive tokens an…

cs.CV2026

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

Ashutosh Kumar, Rajat Saini, Jingjing Pan +5

Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only supervision. We introduce Ins…

cs.CV2025

The 9th AI City Challenge

Zheng Tang, Shuo Wang, David C. Anastasiu +25

The ninth AI City Challenge continues to advance real-world applications of computer vision and AI in transportation, industrial automation, and public safety. The 2025 edition fea…

cs.HC2025

An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance

Hsuan-Kung Yang, Tsu-Ching Hsiao, Ryoichiro Oka +3

Delivering intelligent and adaptive navigation assistance in augmented reality (AR) requires more than visual cues, as it demands systems capable of interpreting flexible user inte…

cs.CV2025

Synthetic Visual Genome

Jae Sung Park, Zixian Ma, Linjie Li +9

Reasoning over visual relationships-spatial, functional, interactional, social, etc.-is considered to be a fundamental component of human cognition. Yet, despite the major advances…

eess.IV2025

Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras

Masataka Kobayashi, Shintaro Shiba, Quan Kong +4

With the growing adoption of autonomous driving, the advancement of sensor technology is crucial for ensuring safety and reliable operation. Sensor fusion techniques that combine m…