8 papers
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
Yuanhao Ban, Tong Xie, Sohyun An +6
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness ben…
A Unifying Lens on Supervised Fine-Tuning Through Target Distribution Design
Tong Xie, Yuanhao Ban, Yunqi Hong +3
Supervised fine-tuning (SFT) typically maximizes the likelihood of every token in a demonstrated trajectory. However, an observed token can be non-unique, noisy, or misaligned with…
FRESCO: Benchmarking and Optimizing Re-rankers for Evolving Semantic Conflict in Retrieval-Augmented Generation
Sohyun An, Hayeon Lee, Shuibenyang Yuan +4
Retrieval-Augmented Generation (RAG) is a key approach to mitigating the temporal staleness of large language models (LLMs) by grounding responses in up-to-date evidence. Within th…
Cycle-Consistent Search: Question Reconstructability as a Proxy Reward for Search Agent Training
Sohyun An, Shuibenyang Yuan, Hayeon Lee +2
Reinforcement Learning (RL) has shown strong potential for optimizing search agents in complex information retrieval tasks. However, existing approaches predominantly rely on gold…
DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
Yu Zhou, Sohyun An, Haikang Deng +5
Contact languages like English exhibit rich regional variations in the form of dialects, which are often used by dialect speakers interacting with generative models. However, can m…
T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search
Hyomin Lee, Sangwoo Park, Yumin Choi +3
While prior red-teaming efforts have focused on eliciting harmful text outputs from large language models (LLMs), such approaches fail to capture agent-specific vulnerabilities tha…