4 papers
Assessing VLM-Driven Semantic-Affordance Inference for Non-Humanoid Robot Morphologies
Jess Jones, Raul Santos-Rodriguez, Sabine Hauert
Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding human-object interactions, but their application to robotic systems with non-humanoid morph…
HiPP-Prune: Hierarchical Preference-Conditioned Structured Pruning for Vision-Language Models
Lincen Bai, Hedi Tabia, Raul Santos-Rodriguez
Pruning vision-language models (VLMs) for efficient deployment is challenging because compression can affect not only task utility but also visual grounding, often amplifying objec…
LEMON: Local Explanations via Modality-aware OptimizatioN
Yu Qin, Phillip Sloan, Raul Santos-Rodriguez +2
Multimodal models are ubiquitous, yet existing explainability methods are often single-modal, architecture-dependent, or too computationally expensive to run at scale. We introduce…
TACTFL: Temporal Contrastive Training for Multi-modal Federated Learning with Similarity-guided Model Aggregation
Guanxiong Sun, Majid Mirmehdi, Zahraa Abdallah +3
Real-world federated learning faces two key challenges: limited access to labelled data and the presence of heterogeneous multi-modal inputs. This paper proposes TACTFL, a unified…