4 papers · 1 filter
Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
Hanhua Hong, Yizhi Li, Jiaoyan Chen +4
Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, where direct paper-to-repositor…
HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
Hanhua Hong, Yizhi LI, Jiaoyan Chen +4
Recent advances in large language models have highlighted their potential to automate computational research, particularly reproducing experimental results. However, existing appro…
Document Reconstruction Unlocks Scalable Long-Context RLVR
Yao Xiao, Lei Wang, Yue Deng +6
Reinforcement Learning with Verifiable Rewards~(RLVR) has become a prominent paradigm to enhance the capabilities (i.e.\ long-context) of Large Language Models~(LLMs). However, it…
Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty
Yao Xiao, Jung-jae Kim, Roy Ka-wei Lee +1
Self-play preference optimization has emerged as a prominent paradigm for aligning large language models (LLMs). It typically involves a language model to generate on-policy respon…