Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
Xinquan Chen, Zhenyun Yin, Shan He +38
As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…
cs.AI2026
Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
Yongyi Wang, Lingfeng Li, Bozhou Chen +5
Recent benchmarks for memory-augmented reinforcement learning (RL) have introduced partially observable Markov decision process (POMDP) environments in which agents must use histor…
cs.AI2026
Decoupling Return-to-Go for Efficient Decision Transformer
Yongyi Wang, Hanyu Liu, Lingfeng Li +5
The Decision Transformer (DT) has established a powerful sequence modeling approach to offline reinforcement learning. It conditions its action predictions on Return-to-Go (RTG), u…