collaborators

7 papers

cs.CV2026

Unified Text-Image Generation with Weakness-Targeted Post-Training

Jiahui Chen, Philippe Hansen-Estruch, Xiaochuang Han +7

Unified multimodal generation architectures that jointly produce text and images have recently emerged as a promising direction for text-to-image (T2I) synthesis. However, many exi…

cs.CL2026

Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image

Yushi Hu, Reyhane Askari-Hemmat, Melissa Hall +3

Reward models (RMs) are essential for training large language models (LLMs), but remain underexplored for omni models that handle interleaved image and text sequences. We introduce…

cs.LG2025

Why Less is More (Sometimes): A Theory of Data Curation

Elvis Dohmatob, Mohammad Pezeshki, Reyhane Askari-Hemmat

This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as clas…

cs.CV2025

Improving the Physics of Video Generation with VJEPA-2 Reward Signal

Jianhao Yuan, Xiaofeng Zhang, Felix Friedrich +7

This is a short technical report describing the winning entry of the PhysicsIQ Challenge, presented at the Perception Test Workshop at ICCV 2025. State-of-the-art video generative…

cs.CV2025

Increasing the Utility of Synthetic Images through Chamfer Guidance

Nicola Dall'Asen, Xiaofeng Zhang, Reyhane Askari Hemmat +4

Conditional image generative models hold considerable promise to produce infinite amounts of synthetic training data. Yet, recent progress in generation quality has come at the exp…

cs.CV2025

Multi-Modal Language Models as Text-to-Image Model Evaluators

Jiahui Chen, Candace Ross, Reyhane Askari-Hemmat +4

The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to…