5 papers · 1 filter
InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
Sarah Rastegar, Violeta Chatalbasheva, Sieger Falkena +5
Text-to-image (T2I) diffusion models generate high-quality images but often fail to capture the spatial relations specified in text prompts. This limitation can be traced to two fa…
Object-Centric Vision Token Pruning for Vision Language Models
Guangyuan Li, Rongzhen Zhao, Jinhong Deng +2
In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning r…
Compositional Scene Understanding through Inverse Generative Modeling
Yanbo Wang, Justin Dauwels, Yilun Du
Generative models have demonstrated remarkable abilities in generating high-fidelity visual content. In this work, we explore how generative models can further be used not only to…
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
Ying Liu, Yijing Hua, Haojiang Chai +2
Open-vocabulary detectors are proposed to locate and recognize objects in novel classes. However, variations in vision-aware language vocabulary data used for open-vocabulary learn…
Compositional Image Decomposition with Diffusion Models
Jocelin Su, Nan Liu, Yanbo Wang +2
Given an image of a natural scene, we are able to quickly decompose it into a set of components such as objects, lighting, shadows, and foreground. We can then envision a scene whe…