agent learning 1large language models 1policy optimization 1reinforcement learning 1transition modeling 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.IR2026
MISO: Model-Internal-State-Guided Optimization for Ranking Models
Yongzhe Zhang, Xiaoyu Deng, Yifan He +29
Ranking models are repeatedly refined within established model families, yet the choice of which component to scale, replace, or retire is often guided by expensive trial-and-error…
cs.LG2026
TAPO: Transition-Aware Policy Optimization for LLM Agents
Cong Li, Peixi Peng, Yisen Zhao +4
The paper introduces TAPO, a training framework that augments reinforcement learning for large language model agents with action‑conditioned next‑observation prediction, improving…