11 papers
Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Johannes Ackermann, Michael Noukhovitch, Takashi Ishida +1
Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Models (LMs). A common problem is reward ha…
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
Issa Sugiura, Daichi Hattori, Kazuo Araragi +5
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important. Unlike existing benchmarks tha…
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
Thanawat Lodkaew, Johannes Ackermann, Soichiro Nishimori +3
A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing de…
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting
Takashi Ishida, Thanawat Lodkaew, Ikko Yamane
Publishing a large language model (LLM) benchmark (especially its ground-truth answers) on the Internet risks contaminating future LLMs and enabling evaluation gaming: it may be un…
Practical estimation of the optimal classification error with soft labels and calibration
Ryota Ushio, Takashi Ishida, Masashi Sugiyama
While the performance of machine learning systems has experienced significant improvement in recent years, relatively little attention has been paid to the fundamental question: to…
Unified Approach for Weakly Supervised Multicalibration
Futoshi Futami, Takashi Ishida
Multicalibration requires predicted scores to agree with label probabilities across rich families of subgroups and score-dependent tests, but existing methods require clean input-l…