activity
20222025
most citedLVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

4 citations · 7 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2024

FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability

Estelle Aflalo, Gabriela Ben Melech Stan, Tiep Le +5

Large Vision Language Models (LVLMs) have achieved significant progress in integrating visual and textual inputs for multimodal reasoning. However, a recurring challenge is ensurin…

cs.CV2024★ 2 cited

L-MAGIC: Language Model Assisted Generation of Images with Coherence

Zhipeng Cai, Matthias Mueller, Reiner Birkl +6

In the current era of generative AI breakthroughs, generating panoramic scenes from a single input image remains a key challenge. Most existing methods use diffusion-based iterativ…

cs.CV2024★ 4 cited

LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Gabriela Ben Melech Stan, Estelle Aflalo, Raanan Yehezkel Rohekar +7

In the rapidly evolving landscape of artificial intelligence, multi-modal large language models are emerging as a significant area of interest. These models, which combine various…

cs.CV2024

Getting it Right: Improving Spatial Consistency in Text-to-Image Models

Agneet Chatterjee, Gabriela Ben Melech Stan, Estelle Aflalo +8

One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in…

cs.CV2022

MuMUR : Multilingual Multimodal Universal Retrieval

Avinash Madasu, Estelle Aflalo, Gabriela Ben Melech Stan +4

Multi-modal retrieval has seen tremendous progress with the development of vision-language models. However, further improving these models require additional labelled data which is…