7 papers
MapTrace: Scalable Data Generation for Route Tracing on Maps
Artemis Panagopoulou, Aveek Purohit, Achin Kulshrestha +2
While Multimodal Large Language Models have achieved human-like performance on many visual and textual reasoning tasks, their proficiency in fine-grained spatial understanding, suc…
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
Artemis Panagopoulou, Le Xue, Honglu Zhou +6
Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressiv…
ViUniT: Visual Unit Tests for More Robust Visual Programming
Artemis Panagopoulou, Honglu Zhou, Silvio Savarese +4
Programming based approaches to reasoning tasks have substantially expanded the types of questions models can answer about visual scenes. Yet on benchmark visual reasoning data, wh…
Evaluating Vision-Language Models on Bistable Images
Artemis Panagopoulou, Coby Melkin, Chris Callison-Burch
Bistable images, also known as ambiguous or reversible images, present visual stimuli that can be seen in two distinct interpretations, though not simultaneously by the observer. I…
X-InstructBLIP: A Framework for aligning X-Modal instruction-aware representations to LLMs and Emergent Cross-modal Reasoning
Artemis Panagopoulou, Le Xue, Ning Yu +7
Recent research has achieved significant advancements in visual reasoning tasks through learning image-to-language projections and leveraging the impressive reasoning abilities of…
Visualizing the Obvious: A Concreteness-based Ensemble Model for Noun Property Prediction
Yue Yang, Artemis Panagopoulou, Marianna Apidianaki +2
Neural language models encode rich knowledge about entities and their relationships which can be extracted from their representations using probing. Common properties of nouns (e.g…