From the 1 of 11 linked papers with an AI index.
11 papers
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré +2
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing trai…
DiMaS: Distribution Matching for Steering Vision-Language-Action Models
Pegah Khayatan, Sara Meziane, Jayneel Parekh +1
The paper introduces DiMaS, a distribution‑matching steering technique that adjusts the internal representations of flow‑matching vision‑language‑action models to achieve fine‑grai…
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
Théophane Vallaeys, Jakob Verbeek, Matthieu Cord
Tokenizers are a key component of state-of-the-art generative image models, extracting the most important features from the signal while reducing data dimension and redundancy. Mos…
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny +3
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the v…
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models
Barbara Toniella Corradini, Mustafa Shukor, Paul Couairon +3
Foundation models have exhibited unprecedented capabilities in tackling many domains and tasks. Models such as CLIP are currently widely used to bridge cross-modal representations,…
Learning to Steer: Input-dependent Steering for Multimodal LLMs
Jayneel Parekh, Pegah Khayatan, Mustafa Shukor +3
Steering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior. However, it remains largely underexplored for multimodal LLM…