Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
Jiaxin Bai, Yue Guo, Yifei Dong +13
World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environment returns text observations after each ac…
cs.CL2026
DeepRefine: Agentic Knowledge Refinement via Reinforcement Learning
Haoyu Huang, Jiaxin Bai, Shujie Liu +7
External knowledge enables large language model (LLM) agents to ground their actions and decisions beyond intrinsic parametric memory in open-ended, knowledge-intensive downstream…