From the 1 of 13 linked papers with an AI index.
1 citations · 1 across the 7 of their papers we have counts for
14 papers
Robust Data-Collection Policy Learning for Low-Variance Online Policy Evaluation
Claire Chen, Shuze Daniel Liu, Licheng Luo +3
In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy s…
Active Offline-to-Online Reinforcement Learning
Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai
The paper proposes an active selection method for fine‑tuning offline‑trained policies under a limited online interaction budget, using upper‑confidence bounds derived from linear…
Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning
Vagul Mahadevan, Claire Chen, Shuze Daniel Liu +1
This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in fast and slow timescales re…
Safe In-Context Reinforcement Learning
Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt +4
In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, in…
Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning
Minjae Kwon, Amir Moeini, Shangtong Zhang +1
Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under…
GameChat: Multi-LLM Dialogue for Safe, Agile, and Socially Optimal Multi-Agent Navigation in Constrained Environments
Vagul Mahadevan, Shangtong Zhang, Rohan Chandra
Safe, agile, and socially compliant multi-robot navigation in cluttered and constrained environments remains a critical challenge. This is especially difficult with self-interested…