8 papers
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
Yipeng Shi, Zhipeng Ma, Yue Wang +4
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimizati…
GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting
Qitai Tan, Ruiwen Gu, Yilin Su +3
Time series forecasting requires models to capture diverse, often mutually exclusive, temporal dynamics, from smooth trend continuation to nonstationary drift and strict phase-alig…
ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
Qitai Tan, Zefang Zong, Yang Li +3
Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teac…
PHGNet: Prototype-Guided Hypergraph Construction for Heterogeneous Spatiotemporal Forecasting
Ruiwen Gu, Yahao Liu, Zhenyu Liu +2
As a core task in intelligent transportation systems, traffic forecasting plays a critical role in urban traffic management. Accurate traffic forecasting relies on modeling complex…
ADMFormer: An Adaptive-Decomposition Transformer with Time-Varying Masked Spatial Attention for Traffic Forecasting
Ruiwen Gu, Qitai Tan, Yahao Liu +1
Accurate traffic forecasting is essential for intelligent transportation systems, supporting a wide range of real-world applications. However, it remains challenging due to two key…
Learning to Commit: Generating Organic Pull Requests via Online Repository Memory
Mo Li, L. H. Xu, Qitai Tan +2
Large language model (LLM)-based coding agents achieve impressive results on controlled benchmarks yet routinely produce pull requests that real maintainers reject. The root cause…