7 papers
Multi-Agent Robotic Control with Onboard Vision-Language Models
Kajetan RachwaÅ, Maciej Majek, BartÅomiej Boczek +6
Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges regarding explainability, general…
Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis
PaweÅ Pozorski, Jakub MuszyÅski, Maria Ganzha
Multimodal Large Language Models (MLLMs) effectively integrate text and audio to interpret context in complex interactive dialogues. However, the internal mechanisms by which heter…
mllm-shap: A Shapley Value Explainability Platform for Text-Audio Multimodal Large Language Models
Jakub MuszyÅski, PaweÅ Pozorski, Maria Ganzha
We introduce mllm-shap, an open-source Python framework designed to extend Shapley Value (SV) explainability from text-only Large Language Models to Multimodal LLMs (MLLMs) process…
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
PaweÅ Pozorski, Jakub MuszyÅski, Maria Ganzha
Explaining the behavior of end-to-end audio language models via Shapley value attribution is intractable under native tokenization: a typical utterance yields over encoder fr…
RAI: Flexible Agent Framework for Embodied AI
Kajetan RachwaÅ, Maciej Majek, BartÅomiej Boczek +4
With an increase in the capabilities of generative language models, a growing interest in embodied AI has followed. This contribution introduces RAI - a framework for creating embo…
Representing and querying data tensors in RDF and SPARQL
Piotr Marciniak, Piotr Sowinski, Maria Ganzha
Embedding tensors in databases has recently gained in significance, due to the rapid proliferation of machine learning methods (including LLMs) which produce embeddings in the form…