2 papers
cs.AI2026
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
Liang He, Jingbo Wen, Hongyu Gu +5
Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval du…
cs.LG2026
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
Liang He, Jingbo Wen, Qishi Zhan +4
Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel. In resource-constrained deployments, the…