3 citations · 5 across the 9 of their papers we have counts for
5 papers · 1 filter
Multimodal Model Diffing for Feature Discovery and Control
Hunar Batra, Lachin Naghashyar, Ashkan Khakzar +4
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control.…
Towards Understanding Multimodal Fine-Tuning: Spatial Features
Lachin Naghashyar, Hunar Batra, Ashkan Khakzar +4
Contemporary Vision-Language Models (VLMs) achieve strong performance on a wide range of tasks by pairing a vision encoder with a pre-trained language model, fine-tuned for visual-…
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
Oishi Deb, Anjun Hu, Ashkan Khakzar +2
We propose a training-free method, Articulate3D, to pose a 3D asset through language control. Despite advances in vision and language models, this task remains surprisingly challen…
Minimalist Concept Erasure in Generative Models
Yang Zhang, Er Jin, Yanfei Dong +5
Recent advances in generative models have demonstrated remarkable capabilities in producing high-quality images, but their reliance on large-scale unlabeled data has raised signifi…
How Visual Representations Map to Language Feature Space in Multimodal LLMs
Constantin Venhoff, Ashkan Khakzar, Sonia Joseph +2
Effective multimodal reasoning depends on the alignment of visual and linguistic representations, yet the mechanisms by which vision-language models (VLMs) achieve this alignment r…