4 papers
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré +2
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing trai…
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny +3
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the v…
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Hugo Caselles-Dupré, Hugo Caselles-Dupré, Mathis Koroglu +3
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4K). Generating videos at the model's nati…
Learning to Steer: Input-dependent Steering for Multimodal LLMs
Jayneel Parekh, Pegah Khayatan, Mustafa Shukor +3
Steering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior. However, it remains largely underexplored for multimodal LLM…