activity
20222026
most citedState Rank Dynamics in Linear Attention LLMs

1 citations · 1 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance

Xinzhu Chen, Wei He, Huichuan Fan +7

Group Relative Policy Optimization (GRPO) performs coarse-grained credit assignment in reinforcement learning with verifiable rewards (RLVR) by assigning the same advantage to all…

cs.CL2026

MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment

Zihao Liu, Hantao Zhou, Jiguo Li +5

User simulators are essential for the scalable training and evaluation of interactive AI systems. However, existing approaches often rely on shallow user profiling, struggle to mai…

cs.CL2026

Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering

Nonghai Zhang, Weitao Ma, Zhanyu Ma +5

Group Relative Policy Optimization (GRPO) significantly enhances the reasoning performance of Large Language Models (LLMs). However, this success heavily relies on expensive extern…

cs.CL2026

UserLM-R1: Modeling Human Reasoning in User Language Models with Multi-Reward Reinforcement Learning

Feng Zhang, Shijia Li, Chunmao Zhang +7

User simulators serve as the critical interactive environment for agent post-training, and an ideal user simulator generalizes across domains and proactively engages in negotiation…

cs.CL2026

Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory Management

Weitao Ma, Xiaocheng Feng, Lei Huang +7

Effective memory management is essential for large language model agents to navigate long-horizon tasks. Recent research has explored using Reinforcement Learning to develop specia…