17 papers
Latent Cluster Analysis for Vision-Language-Action Models
Theodor Wulff, Sergio Lanza, Tamara Bila +3
Vision-Language-Action (VLA) Models are increasingly used in robotics for their ability to ground language and perception into action, yet the internal representations driving thei…
Vision-Language Models are Fragile Multilingual Associators
Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote +2
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes i…
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
Manith Adikari, Bei Peng, Samuele Vinanzi +1
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, c…
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
Theodor Wulff, Federico Tavella, Rahul Singh Maharjan +2
Achieving robot transparency is a critical step toward effective human-robot collaboration. To be transparent, a robot's natural language communication must be consistent with its…
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
Haodong Xie, Yujun Cai, Rahul Singh Maharjan +3
Concept Bottleneck Models (CBMs) introduce interpretability to black-box deep learning models by predicting labels through human-understandable concepts. However, unlike humans, wh…
The Safety Challenge of World Models for Embodied AI Agents: A Review
Lorenzo Baraldi, Zifan Zeng, Chongzhe Zhang +8
The rapid progress in embodied artificial intelligence has highlighted the necessity for more advanced and integrated models that can perceive, interpret, and predict environmental…