From the 1 of 13 linked papers with an AI index.
13 papers
Active Offline-to-Online Reinforcement Learning
Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
The paper proposes an active selection method for fine‑tuning offline‑trained policies under a limited online interaction budget, using upper‑confidence bounds derived from linear…
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
Vagul Mahadevan, Claire Chen, Shuze Daniel Liu +1
This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in fast and slow timescales re…
Safe In-Context Reinforcement Learning
Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt +4
In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, in…
Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning
Minjae Kwon, Amir Moeini, Shangtong Zhang +1
Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under…
GameChat: Multi-LLM Dialogue for Safe, Agile, and Socially Optimal Multi-Agent Navigation in Constrained Environments
Vagul Mahadevan, Shangtong Zhang, Rohan Chandra
Safe, agile, and socially compliant multi-robot navigation in cluttered and constrained environments remains a critical challenge. This is especially difficult with self-interested…
Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning
Alper Kamil Bozkurt, Xiaoan Xu, Shangtong Zhang +2
In offline-to-online reinforcement learning (O2O-RL), policies are first safely trained offline using previously collected datasets and then further fine-tuned for tasks via limite…