5 papers
From Relative Entropy to Minimax: A Unified Framework for Coverage in MDPs
Xihe Gu, Urbashi Mitra, Tara Javidi
Targeted and deliberate exploration of state--action pairs is essential in reward-free Markov Decision Problems (MDPs). More precisely, different state-action pairs exhibit differe…
Partially Decentralized Multi-Agent Q-Learning via Digital Cousins for Wireless Networks
Talha Bozkus, Urbashi Mitra
Q-learning is a widely used reinforcement learning (RL) algorithm for optimizing wireless networks, but faces challenges with large state-spaces. Recently proposed multi-environmen…
Asymmetric Graph Error Control with Low Complexity in Causal Bandits
Chen Peng, Di Zhang, Urbashi Mitra
In this paper, the causal bandit problem is investigated, with the objective of maximizing the long-term reward by selecting an optimal sequence of interventions on nodes in an unk…
A Multi-Agent Multi-Environment Mixed Q-Learning for Partially Decentralized Wireless Network Optimization
Talha Bozkus, Urbashi Mitra
Q-learning is a powerful tool for network control and policy optimization in wireless networks, but it struggles with large state spaces. Recent advancements, like multi-environmen…
Coverage Analysis for Digital Cousin Selection -- Improving Multi-Environment Q-Learning
Talha Bozkus, Tara Javidi, Urbashi Mitra
Q-learning is widely employed for optimizing various large-dimensional networks with unknown system dynamics. Recent advancements include multi-environment mixed Q-learning (MEMQ)…