agent learning 1large language models 1policy optimization 1reinforcement learning 1transition modeling 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
TAPO: Transition-Aware Policy Optimization for LLM Agents
Cong Li, Peixi Peng, Yisen Zhao +4
The paper introduces TAPO, a training framework that augments reinforcement learning for large language model agents with action‑conditioned next‑observation prediction, improving…
cs.RO2025
Delta-Triplane Transformers as Occupancy World Models
Haoran Xu, Peixi Peng, Guang Tan +3
Occupancy World Models (OWMs) aim to predict future scenes via 3D voxelized representations of the environment to support intelligent motion planning. Existing approaches typically…