4 papers · 1 filter
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Chengyu Shen, Yujie Fu, Gangtao Xin +13
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking external tools, and completing complex tas…
One-Eval: An Agentic System for Automated and Traceable LLM Evaluation
Chengyu Shen, Yanheng Hou, Minghui Pan +8
Reliable evaluation is essential for developing and deploying large language models, yet in practice it often requires substantial manual effort: practitioners must identify approp…
DARO: Difficulty-Aware Reweighting Policy Optimization
Jingyu Zhou, Lu Ma, Hao Liang +3
Recent advances in large language models (LLMs) have shown that reasoning ability can be significantly enhanced through Reinforcement Learning with Verifiable Rewards (RLVR). Group…
Let's Verify Math Questions Step by Step
Chengyu Shen, Zhen Hao Wong, Runming He +8
Large Language Models (LLMs) have recently achieved remarkable progress in mathematical reasoning. To enable such capabilities, many existing works distill strong reasoning models…