From the 1 of 5 linked papers with an AI index.
5 papers
DiMaS: Distribution Matching for Steering Vision-Language-Action Models
Pegah Khayatan, Sara Meziane, Jayneel Parekh +1
The paper introduces DiMaS, a distribution‑matching steering technique that adjusts the internal representations of flow‑matching vision‑language‑action models to achieve fine‑grai…
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny +3
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the v…
Learning to Steer: Input-dependent Steering for Multimodal LLMs
Jayneel Parekh, Pegah Khayatan, Mustafa Shukor +3
Steering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior. However, it remains largely underexplored for multimodal LLM…
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
Pegah Khayatan, Mustafa Shukor, Jayneel Parekh +2
Multimodal LLMs (MLLMs) have reached remarkable levels of proficiency in understanding multimodal inputs. However, understanding and interpreting the behavior of such complex model…
A Concept-Based Explainability Framework for Large Multimodal Models
Jayneel Parekh, Pegah Khayatan, Mustafa Shukor +2
Large multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of t…