3 papers
cs.LG2026
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models
Hoang Phan, Xianjun Yang, Yuanshun Yao +6
Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has become a standard post-training paradigm for c…
cs.AI2026
Learning Personalized Agents from Human Feedback
Kaiqu Liang, Julia Kruk, Shengyi Qian +9
Modern AI agents are powerful but often fail to align with the idiosyncratic, evolving preferences of individual users. Prior approaches typically rely on static datasets, either t…
cs.CL2025
Your thoughts tell who you are: Characterize the reasoning patterns of LRMs
Yida Chen, Yuning Mao, Xianjun Yang +7
Current comparisons of large reasoning models (LRMs) focus on macro-level statistics such as task accuracy or reasoning length. Whether different LRMs reason differently remains an…