activity
20242026
collaborators

36 papers

cs.CL2026

Kwai Summary Attention Technical Report

Chenglong Chu, Guorui Zhou, Guowang Zhang +35

Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in semantic understanding/reasoning, code agen…

cs.IR2026

From Extraction to Navigation: Progressive Retrieval with Indirectly Infinite Depth

Linxiao Che, Shanshan Huang, Haitao Lu +6

Modern large-scale recommender retrieval is shifting from static similarity matching to dynamic item space navigation, framing retrieval as iterative goal-driven graph traversal. C…

cs.IR2026

POEM: Partial-Order Enhanced Real-Time Sequential Modeling for Recommendation

Linxiao Che, Yijia Sun, Siyuan Lou +5

Real-time recommendation systems suffer from the dynamic drift of user interests and varying contextual conditions. Conventional sequential recommendation models only exploit stati…

cs.CL2026

When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL

Jiakang Wang, Runze Liu, Qingpeng Cai +7

Reinforcement learning (RL) has shown great promise in large language models (LLMs) post-training, which typically rely on token-level clipping to maintain stability during optimiz…

cs.LG2026

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning

Zhenpeng Su, Leiyu Pan, Minxuan Lv +7

Large language model post-training relies on reinforcement learning to improve model capability and alignment quality. However, the off-policy training paradigm introduces distribu…

cs.LG2026

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning

Zhenpeng Su, Leiyu Pan, Minxuan Lv +5

Reinforcement learning (RL) has become a powerful paradigm for optimizing large language models (LLMs) to handle complex reasoning tasks. A core challenge in this process lies in m…