Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
SkillForge: Evolving Verifiable Skills for Reinforcement Learning Agents
Shidong Yang, Ziyu Ma, Tongwen Huang +5
Large language model (LLM) agents are trained with reinforcement learning (RL) for complex decision-making tasks. However, most RL-trained agents remain episodic and cannot accumul…
cs.CL2026
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding
Hao Zhang, Yiming Hu, Yong Wang +3
Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the target model, enabling lossless…