14 papers
Vision-Language Models are Fragile Multilingual Associators
Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote +2
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable when the language of the input changes i…
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
Manith Adikari, Bei Peng, Samuele Vinanzi +1
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decision-making tasks often involve multiple, c…
LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes
Haozhuo Zhang, Jingkai Sun, Michele Caprio +5
Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism. Almost always, though, this happens in short, isolat…
Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
Federico Tavella, Amber Drinkwater, Angelo Cangelosi
Robotic scene understanding increasingly relies on Vision-Language Models (VLMs) to generate natural language descriptions of the environment. In this work, we systematically evalu…
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
Theodor Wulff, Federico Tavella, Rahul Singh Maharjan +2
Achieving robot transparency is a critical step toward effective human-robot collaboration. To be transparent, a robot's natural language communication must be consistent with its…
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
Haodong Xie, Yujun Cai, Rahul Singh Maharjan +3
Concept Bottleneck Models (CBMs) introduce interpretability to black-box deep learning models by predicting labels through human-understandable concepts. However, unlike humans, wh…