activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation

Yang Zhou, Can Jin, Zihan Dong +7

Reinforcement learning improves the reasoning ability of large language models but remains costly and sample-inefficient, as many rollouts provide weak learning signals. Difficulty…

cs.LG2026

EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning

Wujiang Xu, Wentian Zhao, Zhenting Wang +6

Training LLM agents in multi-turn environments with sparse rewards, where completing a single task requires 30+ turns of interaction within an episode, presents a fundamental chall…

cs.LG2026

AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning

Can Jin, Yang Zhou, Qixin Zhang +8

Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or search-based methods guided by static Proc…

cs.LG2026

Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals

Zihan Dong, Zhixian Zhang, Yang Zhou +3

Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable ranking…

cs.LG2024

Learning from Teaching Regularization: Generalizable Correlations Should be Easy to Imitate

Can Jin, Tong Che, Hongwu Peng +3

Generalization remains a central challenge in machine learning. In this work, we propose Learning from Teaching (LoT), a novel regularization technique for deep neural networks to…