3 papers
cs.CL2026
SafeRun: Enabling Determinism in LLM Planning for Running
Meilin Chen, Zepeng Zhai, Jiaxuan Zhao +1
Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic nature. This limitation is especial…
cs.LG2026
Rewards as Labels: Revisiting RLVR from a Classification Perspective
Zepeng Zhai, Meilin Chen, Jiaxuan Zhao +3
Reinforcement Learning with Verifiable Rewards has recently advanced the capabilities of Large Language Models in complex reasoning tasks by providing explicit rule-based supervisi…
cs.AI2025
Unbiased Evaluation of Large Language Models from a Causal Perspective
Meilin Chen, Jian Tian, Liang Ma +3
Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation o…