3 papers
cs.RO2026
CounterAlign: Counterfactual Supervision for Vision-Language-Action Models
Haru Kondoh, Kei Ota, Asako Kanezaki +1
Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, wi…
cs.CV2025
Embodied Navigation with Auxiliary Task of Action Description Prediction
Haru Kondoh, Asako Kanezaki
The field of multimodal robot navigation in indoor environments has garnered significant attention in recent years. However, as tasks and methods become more advanced, the action d…
cs.CV2023
Multi-goal Audio-visual Navigation using Sound Direction Map
Haru Kondoh, Asako Kanezaki
Over the past few years, there has been a great deal of research on navigation tasks in indoor environments using deep reinforcement learning agents. Most of these tasks use only v…