7 papers
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
Haoxiang Shi, Xiang Deng, Zaijing Li +3
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural language instructions through free-form 3D spaces. Existing VLN-CE approaches typic…
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our met…
OSGNet @ Ego4D Episodic Memory Challenge 2025
Yisen Feng, Haoyu Zhang, Qiaohui Chu +4
In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise…
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
Haoyu Zhang, Meng Liu, Zaijing Li +4
Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied in…
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
Haoyu Zhang, Yisen Feng, Qiaohui Chu +4
In this report, we present the method that achieves third place for Ego4D EgoSchema Challenge in CVPR 2025. To improve the reliability of answer prediction in egocentric video ques…
Efficient Safety Alignment of Large Language Models via Preference Re-ranking and Representation-based Reward Modeling
Qiyuan Deng, Xuefeng Bai, Kehai Chen +3
Reinforcement Learning (RL) algorithms for safety alignment of Large Language Models (LLMs), such as Direct Preference Optimization (DPO), encounter the challenge of distribution s…