2 papers
cs.CV2026
Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models
Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim +1
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spat…
cs.CV2026
Universal Few-Shot Spatial Control for Diffusion Models
Kiet T. Nguyen, Chanhyuk Lee, Donggyun Kim +2
Spatial conditioning in pretrained text-to-image diffusion models has significantly improved fine-grained control over the structure of generated images. However, existing control…