5 papers
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
Yuanhao Ban, Tong Xie, Sohyun An +6
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness ben…
A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design
Tong Xie, Yuanhao Ban, Yunqi Hong +3
Supervised fine-tuning (SFT) typically maximizes the likelihood of every token in a demonstrated trajectory. However, an observed token can be non-unique, noisy, or misaligned with…
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
Zihan Chen, Yiming Zhang, Hengguang Zhou +3
Current benchmarks are inadequate for evaluating progress in reinforcement learning (RL) for large language models (LLMs).Despite recent benchmark gains reported for RL, we find th…
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
Yunqi Hong, Kuei-Chun Kao, Hengguang Zhou +1
Reinforcement learning (RL) has become a standard approach for post-training large language models and, more recently, for improving image generation models, which uses reward func…
Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
Yunqi Hong, Johnson Kao, Liam Edwards +5
AI tools in pathology have improved screening throughput, standardized quantification, and revealed prognostic patterns that inform treatment. However, adoption remains limited bec…