Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning
Yu Li, Shu Hong, Tian Lan
Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecified in long reasoning traces. On-policy sel…
cs.LG2026
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
Yu Li, Rui Miao, Tian Lan +1
Reinforcement learning with verifiable rewards has become the standard recipe for improving LLM reasoning, but the dominant algorithm GRPO assigns a single trajectory-level advanta…
cs.LG2026
ACDZero: MCTS Agent for Mastering Automated Cyber Defense
Yu Li, Sizhe Tang, Rongqian Chen +5
Automated cyber defense (ACD) seeks to protect computer networks with minimal or no human intervention, reacting to intrusions by taking corrective actions such as isolating hosts,…