activity
20242026
collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Diverse-Intent Multi-Turn Fashion Image Retrieval

Mingqiang Tang, Haokun Wen, Meng Liu +3

Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every inte…

cs.CV2026

SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion

Haoyu Zhang, Meng Liu, Qianlong Xiang +3

Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-l…

cs.CV2026

JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026

Qiaohui Chu, Haoyu Zhang, Yisen Feng +4

We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future pred…

cs.CV2026

VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026

Qiaohui Chu, Haoyu Zhang, Yisen Feng +4

We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric vi…

cs.CV2026

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

Yisen Feng, Leigang Qu, Haoyu Zhang +5

In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…

cs.CV2026

MARS: Technical Report for the CASTLE Challenge at EgoVis 2026

Haoyu Zhang, Qiaohui Chu, Yisen Feng +4

This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed-for…