18 papers
The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL
Nicolas Beltran-Velez, Felix Friedrich, Zhang Xiaofeng +4
Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering propertie…
Grounding Computer Use Agents on Human Demonstrations
Aarash Feizi, Shravan Nayak, Xiangru Jian +14
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web…
GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation
Ye Zhu, Kaleb S. Newman, Johannes F. Lutzeyer +3
Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I dive…
PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs
Rim Assouel, Amir Bar, Michal Drozdzal +1
Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this work, we propose Procedurally Ge…
Inference-time Physics Alignment of Video Generative Models with Latent World Models
Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich +7
State-of-the-art video generative models produce promising visual content yet often violate basic physics principles, limiting their utility. While some attribute this deficiency t…
The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
Zhang Xiaofeng, Aaron Courville, Michal Drozdzal +1
Text-to-image (T2I) models offer great potential for creating virtually limitless synthetic data, a valuable resource compared to fixed and finite real datasets. Previous works eva…