8 papers
The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms
Jinghan Zhang, Zerui Cheng, Shiqi Chen +5
Traditional evaluations measure a learning algorithm's final performance on an i.i.d. test set, reducing learning to a single aggregate score. This approach obscures a fundamental…
Learn Hard Problems During RL with Reference Guided Fine-tuning
Yangzhen Wu, Shanda Li, Zixin Wen +5
Reinforcement learning (RL) for mathematical reasoning can suffer from reward sparsity: for challenging problems, LLM fails to sample any correct trajectories, preventing RL from r…
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
Zexuan Wang, Chenghao Yang, Yingqi Que +18
Real-world autonomous planning requires coordinating tightly coupled constraints where a single decision dictates the feasibility of all subsequent actions. However, existing bench…
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
Jiayun Wu, Jiashuo Liu, Zhiyuan Zeng +3
LLM deployment in critical domains is currently impeded by persistent hallucinations--generating plausible but factually incorrect assertions. While scaling laws drove significant…
TabularMath: Evaluating Computational Extrapolation in Tabular Learning via Program-Verified Synthesis
Zerui Cheng, Jiashuo Liu, Jianzhu Yao +3
Standard tabular benchmarks mainly focus on the evaluation of a model's capability to interpolate values inside a data manifold, where models good at performing local statistical s…
VeRA: Verified Reasoning Data Augmentation at Scale
Zerui Cheng, Jiashuo Liu, Chunjie Wu +4
The main issue with most evaluation schemes today is their "static" nature: the same problems are reused repeatedly, allowing for memorization, format exploitation, and eventual sa…