collaborators

9 papers

cs.CV2026

Multimodal Model Diffing for Feature Discovery and Control

Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control.…

cs.CV2026

Towards Understanding Multimodal Fine-Tuning: Spatial Features

Lachin Naghashyar, Hunar Batra, Ashkan Khakzar +4

Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-…

cs.LG2025

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2

Training vision language models (VLMs) aims to align visual representations from a vision encoder with the textual representations of a pretrained large language model (LLM). Howev…

cs.LG2025

RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching

Farnoush Rezaei Jafari, Oliver Eberle, Ashkan Khakzar +1

Activation patching is a standard method in mechanistic interpretability for localizing the components of a model responsible for specific behaviors, but it is computationally expe…

cs.CV2025

Articulate3D: Zero-Shot Text-Driven 3D Object Posing

Oishi Deb, Anjun Hu, Ashkan Khakzar +2

We propose a training-free method, Articulate3D, to pose a 3D asset through language control. Despite advances in vision and language models, this task remains surprisingly challen…

cs.CV2025

Minimalist Concept Erasure in Generative Models

Yang Zhang, Er Jin, Yanfei Dong +5

Recent advances in generative models have demonstrated remarkable capabilities in producing high-quality images, but their reliance on large-scale unlabeled data has raised signifi…