4 papers
SceneForge: Structured World Supervision from 3D Interventions
Jizhizi Li, Jiayang Ao, Danny Wicks +1
Many multimodal learning tasks require supervision that remains consistent across edits, viewpoints, and scene-level interventions. However, such supervision is difficult to obtain…
A Practical Investigation of Spatially-Controlled Image Generation with Transformers
Guoxuan Xia, Harleen Hanspal, Petru-Daniel Tudosiu +2
Enabling image generation models to be spatially controlled is an important area of research, empowering users to better generate images according to their own fine-grained specifi…
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
Raman Dutt, Harleen Hanspal, Guoxuan Xia +5
In this work, we undertake the challenge of augmenting the existing generative capabilities of pre-trained text-only large language models (LLMs) with multi-modal generation capabi…
Generating Compositional Scenes via Text-to-image RGBA Instance Generation
Alessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang +2
Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditio…