2 papers
cs.AI2026
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
Chunji Lv, Yangguang Wei, Junlin Liu +6
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-tur…
cs.LG2026
FuseSampleAgg: One-Pass Neighborhood Estimation for Budgeted Knowledge-Graph Refresh and Validation
Aleksandar StankoviÄ, Haoran Du, Xinming Wang
Operational knowledge-graph (KG) pipelines in networking and cybersecurity increasingly need to refresh embeddings under strict time, memory, and audit budgets, especially as curat…