1 paper
Gershom Seneviratne, Yohan Abeysinghe, Jianyu An +3
We introduce VEGA, an approach for training navigation VisionLanguage-Action (VLA) models from unlabeled egocentric navigation videos. Internet-scale egocentric videos provide a sc…