16 papers · 1 filter
OpenCompass: A Universal Evaluation Platform for Large Language Models
Maosong Cao, Kai Chen, Haodong Duan +27
In recent years, the field of artificial intelligence has undergone a paradigm shift from task-specific small-scale models to general-purpose large language models (LLMs). With the…
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
Xiaoyuan Li, Keqin Bao, Moxin Li +5
Rubric-based rewards offer a promising way to extend reinforcement learning (RL) for large language models beyond tasks with automatically verifiable answers. However, scaling rubr…
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
Xiaoyuan Li, Keqin Bao, Yubo Ma +6
Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly focus on single-turn reasoning s…
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
Xiaoyuan Li, Moxin Li, Keqin Bao +4
Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entries and retrieve them only b…
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
Xiaoyuan Li, Yuzhe Wang, Moxin Li +6
Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabilities remain brittle under qu…
On Predicting the Post-training Potential of Pre-trained LLMs
Xiaoyuan Li, Yubo Ma, Kexin Yang +5
The performance of Large Language Models (LLMs) on downstream tasks is fundamentally constrained by the capabilities acquired during pre-training. However, traditional benchmarks l…