On Myopic Sensing for Multi-Channel Opportunistic Access: Structure, Optimality, and Performance
arXiv:0712.0035 · doi:10.1109/T-WC.2008.071349
Abstract
We consider a multi-channel opportunistic communication system where the states of these channels evolve as independent and statistically identical Markov chains (the Gilbert-Elliot channel model). A user chooses one channel to sense and access in each slot and collects a reward determined by the state of the chosen channel. The problem is to design a sensing policy for channel selection to maximize the average reward, which can be formulated as a multi-arm restless bandit process. In this paper, we study the structure, optimality, and performance of the myopic sensing policy. We show that the myopic sensing policy has a simple robust structure that reduces channel selection to a round-robin procedure and obviates the need for knowing the channel transition probabilities. The optimality of this simple policy is established for the two-channel case and conjectured for the general case based on numerical results. The performance of the myopic sensing policy is analyzed, which, based on the optimality of myopic sensing, characterizes the maximum throughput of a multi-channel opportunistic communication system and its scaling behavior with respect to the number of channels. These results apply to cognitive radio networks, opportunistic transmission in fading environments, and resource-constrained jamming and anti-jamming.
To appear in IEEE Transactions on Wireless Communications. This is a revised version
References in corpus (1)
Cited by in corpus (41)
- Thirty Years of Machine Learning: The Road to Pareto-Optimal Wireless Networks
- On Optimality of Myopic Policy for Restless Multi-armed Bandit Problem with Non i.i.d. Arms and Imperfect Detection
- Cognitive Radar Using Reinforcement Learning in Automotive Applications
- Decentralized Automotive Radar Spectrum Allocation to Avoid Mutual Interference Using Reinforcement Learning
- Learning-Based Multi-Channel Access in 5G and Beyond Networks with Fast Time-Varying Channels
- Deep Reinforcement Learning for Dynamic Multichannel Access in Wireless Networks
- Indexability of Restless Bandit Problems and Optimality of Whittle's Index for Dynamic Multichannel Access
- Applications of Deep Reinforcement Learning in Communications and Networking: A Survey
- Deep Learning based Wireless Resource Allocation with Application to Vehicular Networks
- Learning in Restless Bandits under Exogenous Global Markov Process
- Partially Observable Minimum-Age Scheduling: The Greedy Policy
- A Deep Actor-Critic Reinforcement Learning Framework for Dynamic Multichannel Access
- Learning in A Changing World: Restless Multi-Armed Bandit with Unknown Dynamics
- Actor-Critic Deep Reinforcement Learning for Dynamic Multichannel Access
- Optimality of Myopic Sensing in Multi-Channel Opportunistic Access
- On the Optimality of Myopic Sensing in Multi-channel Opportunistic Access: the Case of Sensing Multiple Channels
- The Non-Bayesian Restless Multi-Armed Bandit: a Case of Near-Logarithmic Regret
- Decentralized Learning for Channel Allocation in IoT Networks over Unlicensed Bandwidth as a Contextual Multi-player Multi-armed Bandit Game
- Security of Spectrum Learning in Cognitive Radios
- Interference Mitigation and Resource Allocation in Underlay Cognitive Radio Networks
- To Stay Or To Switch: Multiuser Dynamic Channel Access
- Exploiting Channel Correlation and PU Traffic Memory for Opportunistic Spectrum Scheduling
- Dynamic Intrusion Detection in Resource-Constrained Cyber Networks
- On Optimality of Myopic Sensing Policy with Imperfect Sensing in Multi-channel Opportunistic Access
- Power Allocation over Two Identical Gilbert-Elliott Channels
- Cost-Aware Learning and Optimization for Opportunistic Spectrum Access
- On Design of Opportunistic Spectrum Access in the Presence of Reactive Primary Users
- Optimal Power Allocation Policy over Two Identical Gilbert-Elliott Channels
- Dynamic Server Allocation over Time Varying Channels with Switchover Delay
- Detecting an Odd Restless Markov Arm with a Trembling Hand
- Sufficient Conditions on the Optimality of Myopic Sensing in Opportunistic Channel Access: A Unifying Framework
- Efficient Online Learning for Opportunistic Spectrum Access
- On The Optimality of Myopic Sensing in Multi-State Channels
- Almost Optimal Channel Access in Multi-Hop Networks With Unknown Channel Variables
- Optimality of Myopic Policy for Restless Multiarmed Bandit with Imperfect Observation
- Delay Sensitive Communications over Cognitive Radio Networks
- Multi-Access Communications with Energy Harvesting: A Multi-Armed Bandit Model and the Optimality of the Myopic Policy
- Algorithms for Dynamic Spectrum Access with Learning for Cognitive Radio
- Delay Optimal Multichannel Opportunistic Access
- Regret Bounds for Opportunistic Channel Access
- Towards Distribution-Free Multi-Armed Bandits with Combinatorial Strategies