works on

From the 1 of 43 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

31 papers · 1 filter

cs.CV2026

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Tianyuan Zhang, Zonglei Jing, Jiangfan Liu +47

The paper reports on the CVPR 2026@AdvML Workshop Challenge, which evaluated adversarial attacks on multimodal vision‑language agents for autonomous driving using multi‑view visual…

cs.CV2026

Training-Free Composed Video Retrieval via Visual Representation-Guided Video-LLM Reasoning

Yang Liu, Qianqian Xu, Peisong Wen +2

Recent advances in large vision-language models have expanded video retrieval from simple text-based search to more flexible scenarios, where users may specify the desired result t…

cs.CV2026

Understanding-Enhanced Model Collaboration for Long-Tailed Egocentric Mistake Detection

Boyu Han, Qianqian Xu, Shilong Bao +3

In this report, we address the problem of determining whether a user performs an action incorrectly from egocentric video data. To this end, we propose an Understanding-Enhanced Mo…

cs.CV2026

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

Zelin Zheng, Xinyan Liu, Ruixin Li +4

Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle num…

cs.CV2026

Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

Zhikang Xu, Qianqian Xu, Zitai Wang +4

Out-of-distribution (OOD) detection seeks to identify samples from unknown classes, a critical capability for deploying machine learning models in open-world scenarios. Recent rese…

cs.CV2026

From Static to Dynamic: Exploring Self-supervised Image-to-Video Representation Transfer Learning

Yang Liu, Qianqian Xu, Peisong Wen +3

Recent studies have made notable progress in video representation learning by transferring image-pretrained models to video tasks, typically with complex temporal modules and video…