From the 1 of 6 linked papers with an AI index.
6 papers
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
Yipeng Shi, Zhipeng Ma, Yue Wang +4
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimizati…
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
Yuanjie Lu, Beichen Wang, Zhengqi Wu +4
The paper introduces APPLV, a system that uses vision‑language models to predict parameters for classical motion planners, combining safety of traditional planners with adaptabilit…
ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks
Qitai Tan, Zefang Zong, Yang Li +3
Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teac…
ATGPO: Agentic Turn-Group Policy Optimization with Adaptive Turn-level Clipping
Dingwei Chen, Zefang Zong, Zhipeng Ma +5
Reinforcement learning for agentic large language models (LLMs) typically relies on a sparse, trajectory-level outcome reward, making it difficult to evaluate the contribution of i…
Task Expansion and Cross Refinement for Open-World Conditional Modeling
Shreyas Bhat Brahmavar, Qiyang Liu, Yang Li +1
Open-world conditional modeling (OCM), requires a single model to answer arbitrary conditional queries across heterogeneous datasets, where observed variables and targets vary and…
Towards Universal Neural Likelihood Inference
Shreyas Bhat Brahmavar, Yang Li, Qiyang Liu +2
We introduce universal neural likelihood inference (UNLI): enabling a single model to provide data-grounded, conditional likelihood predictions for arbitrary targets given any coll…