most citedState Rank Dynamics in Linear Attention LLMs

1 citations · 1 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CL2026

Hidden States Know Where Reasoning Diverges: Credit Assignment via Span-Level Wasserstein Distance

Xinzhu Chen, Wei He, Huichuan Fan +7

Group Relative Policy Optimization (GRPO) performs coarse-grained credit assignment in reinforcement learning with verifiable rewards (RLVR) by assigning the same advantage to all…

cs.CL2026

MUSE: Multi-Domain Chinese User Simulation via Self-Evolving Profiles and Rubric-Guided Alignment

Zihao Liu, Hantao Zhou, Jiguo Li +5

User simulators are essential for the scalable training and evaluation of interactive AI systems. However, existing approaches often rely on shallow user profiling, struggle to mai…

cs.CL2026

Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering

Nonghai Zhang, Weitao Ma, Zhanyu Ma +5

Group Relative Policy Optimization (GRPO) significantly enhances the reasoning performance of Large Language Models (LLMs). However, this success heavily relies on expensive extern…

cs.LG20261 cited

State Rank Dynamics in Linear Attention LLMs

Ao Sun, Hongtao Zhang, Heng Zhou +9

Linear Attention Large Language Models (LLMs) offer a compelling recurrent formulation that compresses context into a fixed-size state matrix, enabling constant-time inference. How…

cs.LG2026

From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement Learning

Wenzhe Niu, Wei He, Zongxia Xie +10

Reinforcement learning has become a cornerstone for enhancing the reasoning capabilities of Large Language Models, where group-based approaches such as GRPO have emerged as efficie…

cs.SD2026

LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning

Wenhao Zou, Yuwei Miao, Zhanyu Ma +5

Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in…