works on

From the 1 of 13 linked papers with an AI index.

collaborators

13 papers

cs.LG2026

Active Offline-to-Online Reinforcement Learning

Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai

The paper proposes an active selection method for fine‑tuning offline‑trained policies under a limited online interaction budget, using upper‑confidence bounds derived from linear…

cs.LG2026

Convergence of Two-Timescale Markovian Stochastic Approximations with Applications in Reinforcement Learning

Vagul Mahadevan, Claire Chen, Shuze Daniel Liu +1

This work studies the convergence of two-timescale stochastic approximations (SA), a class of iterative algorithms that update two sets of parameters in fast and slow timescales re…

cs.LG2026

Safe In-Context Reinforcement Learning

Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt +4

In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, in…

cs.LG2026

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

Minjae Kwon, Amir Moeini, Shangtong Zhang +1

Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cost under a safety budget. Under…

cs.RO2026

GameChat: Multi-LLM Dialogue for Safe, Agile, and Socially Optimal Multi-Agent Navigation in Constrained Environments

Vagul Mahadevan, Shangtong Zhang, Rohan Chandra

Safe, agile, and socially compliant multi-robot navigation in cluttered and constrained environments remains a critical challenge. This is especially difficult with self-interested…

cs.LG2026

Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning

Alper Kamil Bozkurt, Xiaoan Xu, Shangtong Zhang +2

In offline-to-online reinforcement learning (O2O-RL), policies are first safely trained offline using previously collected datasets and then further fine-tuned for tasks via limite…