Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models
Zongji Yu, Wenshui Luo, Yiliu Sun +4
Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimizat…
cs.CL2026
KALE: Enhancing Knowledge Manipulation in Large Language Models via Knowledge-aware Learning
Qitan Lv, Tianyu Liu, Qiaosheng Zhang +2
Despite the impressive performance of large language models (LLMs) pretrained on vast knowledge corpora, advancing their knowledge manipulation-the ability to effectively recall, r…