2 papers
cs.RO2026
CounterAlign: Counterfactual Supervision for Vision-Language-Action Models
Haru Kondoh, Kei Ota, Asako Kanezaki +1
Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, wi…
cs.CV2025
Embodied Navigation with Auxiliary Task of Action Description Prediction
Haru Kondoh, Asako Kanezaki
The field of multimodal robot navigation in indoor environments has garnered significant attention in recent years. However, as tasks and methods become more advanced, the action d…