3 papers
cs.AI2026
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
Jinkun Hou, Zhuo Liu, Huimin Ren +3
Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without following a single correct generation tra…
cs.LG2026
TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition
Ningyuan Xi, Hao Xu, Hongsheng Xin +1
Large language models (LLMs) have made remarkable progress in reasoning tasks, largely driven by post-training paradigms, especially reinforcement learning with verifiable rewards…
cs.CL2026
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions
Xuan Yang, Hao Xu, Tingfeng Hui +4
Despite great advances in tool-use capabilities of large language models (LLMs), existing evaluation benchmarks struggle to fully align with real-world scenarios. Such benchmarks m…