Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
PRBench: End-to-end Paper Reproduction in Physics Research
Shi Qiu, Junyi Deng, Yiwei Deng +48
AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation a…
cs.CL2025
DIVE: Diversified Iterative Self-Improvement
Yiwei Qin, Yixiu Liu, Pengfei Liu
Recent advances in large language models (LLMs) have demonstrated the effectiveness of Iterative Self-Improvement (ISI) techniques. However, continuous training on self-generated d…
cs.CL2024★ 1 cited
InFoBench: Evaluating Instruction Following Ability in Large Language Models
Yiwei Qin, Kaiqiang Song, Yebowen Hu +7
This paper introduces the Decomposed Requirements Following Ratio (DRFR), a new metric for evaluating Large Language Models' (LLMs) ability to follow instructions. Addressing a gap…