works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CL2026

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

Hanhua Hong, Yizhi Li, Jiaoyan Chen +4

The paper conducts a meta‑evaluation of rubrics generated by large language models for assessing the reproducibility of research papers, comparing intrinsic semantic similarity and…

cs.AI2026

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

Chengwei Wei, Jung-jae Kim, Longyin Zhang +2

Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary computational cost. While existing re…

cs.CL2026

HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution

Hanhua Hong, Yizhi LI, Jiaoyan Chen +4

Recent advances in large language models have highlighted their potential to automate computational research, particularly reproducing experimental results. However, existing appro…

cs.CL2026

Document Reconstruction Unlocks Scalable Long-Context RLVR

Yao Xiao, Lei Wang, Yue Deng +6

Reinforcement Learning with Verifiable Rewards~(RLVR) has become a prominent paradigm to enhance the capabilities (i.e.\ long-context) of Large Language Models~(LLMs). However, it…

cs.CL2026

Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty

Yao Xiao, Jung-jae Kim, Roy Ka-wei Lee +1

Self-play preference optimization has emerged as a prominent paradigm for aligning large language models (LLMs). It typically involves a language model to generate on-policy respon…