works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization

Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré +2

Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing trai…

cs.CV2026

SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization

Théophane Vallaeys, Jakob Verbeek, Matthieu Cord

Tokenizers are a key component of state-of-the-art generative image models, extracting the most important features from the signal while reducing data dimension and redundancy. Mos…

cs.CV2026

When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs

Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny +3

Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the v…

cs.CV2025

FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models

Barbara Toniella Corradini, Mustafa Shukor, Paul Couairon +3

Foundation models have exhibited unprecedented capabilities in tackling many domains and tasks. Models such as CLIP are currently widely used to bridge cross-modal representations,…

cs.CV2025

OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models

Mathis Koroglu, Hugo Caselles-Dupré, Guillaume Jeanneret Sanmiguel +1

We consider the problem of text-to-video generation tasks with precise control for various applications such as camera movement control and video-to-video editing. Most methods tac…

cs.CV2025

DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut

Paul Couairon, Mustafa Shukor, Jean-Emmanuel Haugeard +2

Foundation models have emerged as powerful tools across various domains including language, vision, and multimodal tasks. While prior works have addressed unsupervised image segmen…