12 papers
Diverse-Intent Multi-Turn Fashion Image Retrieval
Mingqiang Tang, Haokun Wen, Meng Liu +3
Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every inte…
SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
Haoyu Zhang, Meng Liu, Qianlong Xiang +3
Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-l…
KAN-MLP-Mixer: A comprehensive investigation of the usage of Kolmogorov-Arnold Networks (KANs) for improving IMU-based Human Activity Recognition
Mengxi Liu, Sizhen Bian, Vitor Fortes +5
Kolmogorov-Arnold Networks (KANs) have demonstrated an exceptional ability to learn complex functions on clean, low-dimensional data but struggle to maintain performance on noisy a…
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future pred…
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric vi…
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
Yisen Feng, Leigang Qu, Haoyu Zhang +5
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…