activity
20242026
collaborators

5 papers

cs.CV2026

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

Martina Ianaro, Guilherme Fernandes, Maurizio Gabbrielli +1

As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perc…

cs.CV2026

Early Estimation of Language to Latent Alignment in Diffusion Models

Vasco Ramos, Regev Cohen, Idan Szpektor +1

Conditional diffusion models frequently suffer from language-image misalignments. Due to the ambiguity of intermediate noise corrupted latents, assessing prompt adherence currently…

cs.CV2025

Latent Beam Diffusion Models for Generating Visual Sequences

Guilherme Fernandes, Vasco Ramos, Regev Cohen +2

While diffusion models excel at generating high-quality images from text prompts, they struggle with visual consistency when generating image sequences. Existing methods generate e…

cs.CV2024

Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis

Vasco Ramos, Yonatan Bitton, Michal Yarom +2

Generated video scenes for action-centric sequence descriptions, such as recipe instructions and do-it-yourself projects, often include non-linear patterns, where the next video ma…

cs.CV2024

Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks

João Bordalo, Vasco Ramos, Rodrigo Valério +5

Multistep instructions, such as recipes and how-to guides, greatly benefit from visual aids, such as a series of images that accompany the instruction steps. While Large Language M…