153 citations · 170 across the 4 of their papers we have counts for
6 papers
Artificial Intelligence and Life in 2030: The One Hundred Year Study on Artificial Intelligence
Peter Stone, Rodney Brooks, Erik Brynjolfsson +14
In September 2016, Stanford's "One Hundred Year Study on Artificial Intelligence" project (AI100) issued the first report of its planned long-term periodic assessment of artificial…
An Analysis of Frame-skipping in Reinforcement Learning
Shivaram Kalyanakrishnan, Siddharth Aravindan, Vishwajeet Bagdawat +5
In the practice of sequential decision making, agents are often designed to sense state at regular intervals of time steps, , ignoring state information in between sensi…
Lower Bounds for Policy Iteration on Multi-action MDPs
Kumar Ashutosh, Sarthak Consul, Bhishma Dedhia +3
Policy Iteration (PI) is a classical family of algorithms to compute an optimal policy for any given Markov Decision Problem (MDP). The basic idea in PI is to begin with some initi…
Regret Minimisation in Multi-Armed Bandits Using Bounded Arm Memory
Arghya Roy Chaudhuri, Shivaram Kalyanakrishnan
In this paper, we propose a constant word (RAM model) algorithm for regret minimisation for both finite and infinite Stochastic Multi-Armed Bandit (MAB) instances. Most of the exis…
PAC Identification of Many Good Arms in Stochastic Multi-Armed Bandits
Arghya Roy Chaudhuri, Shivaram Kalyanakrishnan
We consider the problem of identifying any out of the best arms in an -armed stochastic multi-armed bandit. Framed in the PAC setting, this particular problem generalise…
RLWS: A Reinforcement Learning based GPU Warp Scheduler
Jayvant Anantpur, Nagendra Gulur Dwarakanath, Shivaram Kalyanakrishnan +2
The Streaming Multiprocessors (SMs) of a Graphics Processing Unit (GPU) execute instructions from a group of consecutive threads, called warps. At each cycle, an SM schedules a war…