Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents
Yihong Tang, Kehai Chen, Liang Yue +2
Recent advancements in Reinforcement Learning (RL), particularly Group Relative Policy Optimization (GRPO), have significantly enhanced the reasoning capabilities of Large Language…
cs.CL2026
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
Dingfeng Jiang, Han Yan, Chenze Ma +12
Medical large language models hold promise for reducing healthcare disparities, yet Hindi remains severely underrepresented. While medical LLMs excel in high-resource languages, th…
cs.CL2026
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
Yihong Tang, Kehai Chen, Xuefeng Bai +4
Current role-playing agents (RPAs) are typically constructed by imitating surface-level behaviors, but this approach lacks internal cognitive consistency, often causing out-of-char…