3 papers
cs.AI2026
Bellman-Taylor Score Decoding for Markov Decision Processes with State-Dependent Feasible Action Sets
Yi Chen, Rushuai Yang, Qiang Chen +2
Many Markov decision processes (MDPs) in operations research have feasible actions that are state dependent and defined implicitly by various operational constraints. These feature…
cs.AI2026
ConMem: Structured Memory-Guided Adaptation in Training-Free Multi-Agent Systems
Zhixun Tan, Qiang Chen, Tairan Huang +2
Recent advances have improved the adaptive capabilities of LLM-based multi-agent systems (MAS) through memory-, skill-, and learning-based approaches, yet these approaches remain c…
cs.CL2025
Supervised Optimism Correction: Be Confident When LLMs Are Sure
Junjie Zhang, Rushuai Yang, Shunyu Liu +5
In this work, we establish a novel theoretical connection between supervised fine-tuning and offline reinforcement learning under the token-level Markov decision process, revealing…