4 papers · 1 filter
Distilling Specialized Orders for Visual Generation
Rishav Pramanik, Amin Sghaier, Masih Aminbeidokhti +6
Autoregressive (AR) image generators are becoming increasingly popular due to their ability to produce high-quality images and their scalability. Typical AR models are locked onto…
SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis
Rishav Pramanik, Ian E. Nielsen, Jeff Smith +3
The rapid progress of text-to-image (T2I) models has unlocked unprecedented creative potential, yet their ability to faithfully render complex prompts involving multiple objects, a…
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Juan A. Rodriguez, Haotian Zhang, Abhay Puri +12
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in vision-language models (VLMs) have enabled high-qua…
Masked Multi-Query Slot Attention for Unsupervised Object Discovery
Rishav Pramanik, José-Fabian Villa-Vásquez, Marco Pedersoli
Unsupervised object discovery is becoming an essential line of research for tackling recognition problems that require decomposing an image into entities, such as semantic segmenta…