6 papers
Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment
Theodor Wulff, Federico Tavella, Rahul Singh Maharjan +2
Achieving robot transparency is a critical step toward effective human-robot collaboration. To be transparent, a robot's natural language communication must be consistent with its…
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
Haodong Xie, Yujun Cai, Rahul Singh Maharjan +3
Concept Bottleneck Models (CBMs) introduce interpretability to black-box deep learning models by predicting labels through human-understandable concepts. However, unlike humans, wh…
Joint Action Language Modelling for Transparent Policy Execution
Theodor Wulff, Rahul Singh Maharjan, Xinyun Chi +1
An agent's intention often remains hidden behind the black-box nature of embodied policies. Communication using natural language statements that describe the next action can provid…
Attributes-aware Visual Emotion Representation Learning
Rahul Singh Maharjan, Marta Romeo, Angelo Cangelosi
Visual emotion analysis or recognition has gained considerable attention due to the growing interest in understanding how images can convey rich semantics and evoke emotions in hum…
From Concrete to Abstract: A Multimodal Generative Approach to Abstract Concept Learning
Haodong Xie, Rahul Singh Maharjan, Federico Tavella +1
Understanding and manipulating concrete and abstract concepts is fundamental to human intelligence. Yet, they remain challenging for artificial agents. This paper introduces a mult…
Noise-Free Explanation for Driving Action Prediction
Hongbo Zhu, Theodor Wulff, Rahul Singh Maharjan +2
Although attention mechanisms have achieved considerable progress in Transformer-based architectures across various Artificial Intelligence (AI) domains, their inner workings remai…