activity
20202025
most citedAn Introduction to Vision-Language Modeling

37 citations · 50 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

SneakPeek: Future-Guided Instructional Streaming Video Generation

Cheeun Hong, German Barquero, Fadime Sener +6

Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad imp…

cs.CV2025

Controlling Multimodal LLMs via Reward-guided Decoding

Oscar Mañas, Pierluca D'Oro, Koustuv Sinha +3

As Multimodal Large Language Models (MLLMs) gain widespread applicability, it is becoming increasingly desirable to adapt them for diverse user needs. In this paper, we study the a…

cs.CV2025

Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection

Shivam Chandhok, Qian Yang, Oscar Manas +3

Instruction tuning has been central to the success of recent vision-language models (VLMs), but it remains expensive-requiring large-scale datasets, high-quality annotations, and l…

cs.CV2024

EvalGIM: A Library for Evaluating Generative Image Models

Melissa Hall, Oscar Mañas, Reyhane Askari-Hemmat +14

As the use of text-to-image generative models increases, so does the adoption of automatic benchmarking methods used in their evaluation. However, while metrics and datasets abound…

cs.CV2024★ 3 cited

Consistency-diversity-realism Pareto fronts of conditional image generative models

Pietro Astolfi, Marlene Careil, Melissa Hall +5

Building world models that accurately and comprehensively represent the real world is the utmost aspiration for conditional image generative models as it would enable their use as…

cs.CV2024★ 6 cited

Improving Text-to-Image Consistency via Automatic Prompt Optimization

Oscar Mañas, Pietro Astolfi, Melissa Hall +6

Impressive advances in text-to-image (T2I) generative models have yielded a plethora of high performing models which are able to generate aesthetically appealing, photorealistic im…