Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
Yang Tian, Rui Wang, Xumeng Wen +5
Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide lim…
cs.LG2026
IntentKV: Cross-Turn Intent-Aware KV Cache Pruning for Agent Inference
Junjie Li, Jiong Lou, Jie Li
Multi-turn LLM agents fan short queries into long trajectories of tool calls, search results, and intermediate reasoning. Both KV memory and KV read bandwidth grow by orders of mag…