6 papers
From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video
Qiaohui Chu, Haoyu Zhang, Meng Liu +3
Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important ca…
Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation
Haoxiang Shi, Xiang Deng, Haoyu Zhang +3
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-like environments. Most VLN-CE…
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future pred…
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric vi…
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
Yisen Feng, Leigang Qu, Haoyu Zhang +5
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
Haoyu Zhang, Qiaohui Chu, Yisen Feng +4
This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed-for…