collaborators

6 papers

cs.CV2026

From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video

Qiaohui Chu, Haoyu Zhang, Meng Liu +3

Egocentric 4D interaction forecasting aims to anticipate both where future interactions will occur in 3D and how the human body will move to realize them, providing an important ca…

cs.RO2026

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation

Haoxiang Shi, Xiang Deng, Haoyu Zhang +3

Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-like environments. Most VLN-CE…

cs.CV2026

JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026

Qiaohui Chu, Haoyu Zhang, Yisen Feng +4

We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future pred…

cs.CV2026

VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026

Qiaohui Chu, Haoyu Zhang, Yisen Feng +4

We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric vi…

cs.CV2026

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

Yisen Feng, Leigang Qu, Haoyu Zhang +5

In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 2026. Both tracks require accu…

cs.CV2026

MARS: Technical Report for the CASTLE Challenge at EgoVis 2026

Haoyu Zhang, Qiaohui Chu, Yisen Feng +4

This report presents MARS, short for Multimodal Agentic Reasoning with Source selection, our system for the CASTLE Challenge at EgoVis 2026. Participants must answer 185 closed-for…