9 papers
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
Niantong Li, Guangzheng Hu, Weixu Qiao +35
Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple text-image alignment can no…
Video-Based Reward Modeling for Computer-Use Agents
Linxin Song, Jieyu Zhang, Huanxin Sheng +6
Computer-using agents (CUAs) are becoming increasingly capable; however, it remains difficult to scale evaluation of whether a trajectory truly fulfills a user instruction. In this…
MusicSem: A Semantically Rich Language--Audio Dataset of Natural Music Descriptions
Rebecca Salganik, Teng Tu, Fei-Yueh Chen +8
Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio…
"Rebuilding" Statistics in the Age of AI: A Town Hall Discussion on Culture, Infrastructure, and Training
David L. Donoho, Jian Kang, Xihong Lin +7
This article presents the full, original record of the 2024 Joint Statistical Meetings (JSM) town hall, "Statistics in the Age of AI," which convened leading statisticians to discu…
CLIMB: Class-imbalanced Learning Benchmark on Tabular Data
Zhining Liu, Zihao Li, Ze Yang +6
Class-imbalanced learning (CIL) on tabular data is important in many real-world applications where the minority class holds the critical but rare outcomes. In this paper, we presen…
Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction
Huanxin Sheng, Xinyi Liu, Hangfeng He +2
LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains u…