Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
Guillaume Jeanneret, Mathis Koroglu, Hugo Caselles-Dupré +2
Diffusion Transformer Text-to-Video models have achieved remarkable synthesis quality, yet fine-grained spatial controllability remains a significant challenge. While existing trai…
cs.CV2026
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny +3
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the v…
cs.CV2026
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Hugo Caselles-Dupré, Hugo Caselles-Dupré, Mathis Koroglu +3
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4K). Generating videos at the model's nati…