collaborators

6 papers

cs.LG2026

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

Guanqun Zhao, Zijun Xie, Binbin Zheng +3

Asynchronous reinforcement learning has become the standard way to scale training for large language models (LLM), but the resulting policy lag biases the critic toward the stale b…

cs.CL2026

Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry

Yehan Yang, Junyuan Shang, Yang Li +3

Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide whic…

cs.AI2026

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

Binbin Zheng, Zijun Xie, Guanqun Zhao +4

Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervis…

cs.AI2026

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

Guanqun Zhao, Zijun Xie, Binbin Zheng +5

Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, o…

cs.LG2026

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL

Zijun Xie, Binbin Zheng, Enlei Gong +7

Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollou…

cs.LG2026

Geometry-Aware Contrastive Learning for Few-Shot Automatic Modulation Recognition

Guanqun Zhao, Yitong Liu, Jiaxuan Fang +2

Standard Self-Supervised Learning (SSL) for Automatic Modulation Recognition (AMR) struggles with ineffective isotropic augmentations, spectral instability, and semantic drift. To…