2 papers
cs.RO2025
Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization
Jonathan Yang, Chuyuan Kelly Fu, Dhruv Shah +3
In this work, we investigate how spatially grounded auxiliary representations can provide both broad, high-level grounding as well as direct, actionable information to improve poli…
cs.RO2024
Vision Language Models are In-Context Value Learners
Yecheng Jason Ma, Joey Hejna, Ayzaan Wahid +15
Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress estimator, or temporal…