19 papers
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Yujin Kim, Namgyu Ho, Sangmin Hwang +7
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt th…
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
Young-Jun Lee, Seungone Kim, Minki Kang +5
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search h…
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
Young-Jun Lee, Seungone Kim, Byung-Kwan Lee +6
Can language models (LMs) self-refine their own responses? This question is increasingly relevant as a wide range of real-world user interactions involve refinement requests. Howev…
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
Hyungjoo Chae, Sunghwan Kim, Junhee Cho +18
Web navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimo…
M-Prometheus: A Suite of Open Multilingual LLM Judges
José Pombal, Dongkeun Yoon, Patrick Fernandes +5
The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English,…
Reasoning Models Better Express Their Confidence
Dongkeun Yoon, Seungone Kim, Sohee Yang +6
Despite their strengths, large language models (LLMs) often fail to communicate their confidence accurately, making it difficult to assess when they might be wrong and limiting the…