5 papers
Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation
Haoxiang Shi, Xiang Deng, Haoyu Zhang +3
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-like environments. Most VLN-CE…
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future pred…
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric vi…
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
Yisen Feng, Leigang Qu, Haoyu Zhang +5
In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
Haoyu Zhang, Qiaohui Chu, Yisen Feng +4
This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed-for…