4 papers
CML-Bench: A Framework for Evaluating and Enhancing LLM-Powered Movie Scripts Generation
Mingzhe Zheng, Dingjie Song, Guanyu Zhou +7
Large Language Models (LLMs) have demonstrated remarkable proficiency in generating highly structured texts. However, while exhibiting a high degree of structural organization, mov…
TreeMind: Automatically Reproducing Android Bug Reports via LLM-empowered Monte Carlo Tree Search
Zhengyu Chen, Zhaoyi Meng, Wenxiang Zhao +5
Automatically reproducing Android app crashes from textual bug reports is challenging, particularly when the reports are incomplete and the modern UI exhibits high combinatorial co…
MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
Mohan Jiang, Jin Gao, Jiahao Zhan +1
As multimodal large language models (MLLMs) grow increasingly capable, fixed benchmarks are gradually losing their effectiveness in evaluating high-level scientific understanding.…
Fake it till You Make it: Reward Modeling as Discriminative Prediction
Runtao Liu, Jiahao Zhan, Yingqing He +3
An effective reward model plays a pivotal role in reinforcement learning for post-training enhancement of visual generative models. However, current approaches of reward modeling s…