12 papers
OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
Mengzhang Cai, Xin Gao, Yu Li +13
The rapid evolution of Large Language Models (LLMs) is predicated on the quality and diversity of post-training datasets. However, a critical dichotomy persists: while models are r…
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
Honglin Lin, Qizhi Pei, Xin Gao +5
Reasoning capability is pivotal for Large Language Models (LLMs) to solve complex tasks, yet achieving reliable and scalable reasoning remains challenging. While Chain-of-Thought (…
ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
Qizhi Pei, Zhuoshi Pan, Honglin Lin +6
Large Reasoning Models (LRMs) have shown impressive capabilities in complex problem-solving, often benefiting from training on difficult mathematical problems that stimulate intric…
REST: Stress Testing Large Reasoning Models by Asking Multiple Problems at Once
Zhuoshi Pan, Qizhi Pei, Yu Li +5
Recent Large Reasoning Models (LRMs) have achieved remarkable progress on task-specific benchmarks, yet their evaluation methods remain constrained by isolated problem-solving para…
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
Chenlin Ming, Chendi Qu, Mengzhang Cai +6
Large Language Models (LLMs) have achieved impressive performance through Supervised Fine-tuning (SFT) on diverse instructional datasets. When training on multiple capabilities sim…
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges
Yu Li, Qizhi Pei, Mengyuan Sun +6
Large language models (LLMs) have demonstrated remarkable capabilities, especially the recent advancements in reasoning, such as o1 and o3, pushing the boundaries of AI. Despite th…