collaborators

7 papers

cs.CL2026

Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM

Xiaomeng Hu, Jiaqi Hu, Hao Chen +4

With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex pro…

cs.AI2026

Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization

Hao Chen, Zhanming Shen, Liyao Li +8

Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for eliciting long-chain reasoning in large language models. However, existing methods base…

cs.IR2026

OneReason Technical Report

OneRec Team, Biao Yang, Boyang Ding +81

Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. Howev…

cs.CL2026

SkillComposer: Learning to Evolve Agent Skills for Specification and Generalization

Qi Zhang, Zhaopeng Feng, Xiaonan Shi +8

Agent skills, which consist of reusable strategies that guide agent reasoning and action, have shown strong potential for improving model capability at inference time. However, cur…

cs.CL2026

DeltaMem: Towards Agentic Memory Management via Reinforcement Learning

Qi Zhang, Shen Huang, Chu Liu +4

Recent advances in persona-centric memory have revealed the powerful capability of multi-agent systems in managing persona memory, especially in conversational scenarios. However,…

cs.CL2025

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization

Qi Zhang, Shouqing Yang, Lirong Gao +8

Large language models (LLMs) have demonstrated impressive capabilities in reasoning with the emergence of reasoning models like OpenAI-o1 and DeepSeek-R1. Recent research focuses o…