3 papers
cs.AI2026
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization
Xiwen Chen, Wenhui Zhu, Jingjing Wang +13
Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO…
cs.CL2026
AriadneMem: Threading the Maze of Lifelong Memory for LLM Agents
Wenhui Zhu, Xiwen Chen, Zhipeng Wang +11
Long-horizon LLM agents require memory systems that remain accurate under fixed context budgets. However, existing systems struggle with two persistent challenges in long-term dial…
cs.CL2025
Learning to Reason with Mixture of Tokens
Adit Jain, Brendan Rappazzo
Reinforcement learning with verifiable rewards (RLVR) has become a leading approach for improving large language model (LLM) reasoning capabilities. Most current methods follow var…