collaborators

5 papers

cs.CV2026

Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus

Waikit Xiu, Qiang Lu, Zian Wang +4

In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language Models (MLLMs) tend to over-a…

cs.CV2026

Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning

Waikit Xiu, Qiang Lu, Bingchen Liu +2

For safe and robust autonomous driving, decision-making systems must effectively leverage past experiences to handle the inherent long-tail of traffic scenarios. Case-Based Reasoni…

cs.CV2025

Spacewalk-18: A Benchmark for Multimodal and Long-form Procedural Video Understanding in Novel Domains

Zitian Tang, Rohan Myer Krishnan, Zhiqiu Yu +1

Learning from (procedural) videos has increasingly served as a pathway for embodied agents to acquire skills from human demonstrations. To do this, video understanding models must…

cs.CV2025

How Can Objects Help Video-Language Understanding?

Zitian Tang, Shijie Wang, Junho Cho +2

Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which obj…

cs.CV2025

GMM-Based Comprehensive Feature Extraction and Relative Distance Preservation For Few-Shot Cross-Modal Retrieval

Chengsong Sun, Weiping Li, Xiang Li +2

Few-shot cross-modal retrieval focuses on learning cross-modal representations with limited training samples, enabling the model to handle unseen classes during inference. Unlike t…