2 papers
cs.LG2026
MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents
Bo Qian, Yuting Wu, Shuang Zeng +3
Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Existing methods refine trajectory-level sig…
cs.AI2024
CognTKE: A Cognitive Temporal Knowledge Extrapolation Framework
Wei Chen, Yuting Wu, Shuhan Wu +4
Reasoning future unknowable facts on temporal knowledge graphs (TKGs) is a challenging task, holding significant academic and practical values for various fields. Existing studies…