activity
20112022
most citedWhen to look at a noisy Markov chain in sequential decision making if measurements are costly?

6 citations · 26 across the 23 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2022

Inverse-Inverse Reinforcement Learning. How to Hide Strategy from an Adversarial Inverse Reinforcement Learner

Kunal Pattanayak, Vikram Krishnamurthy, Christopher Berry

Inverse reinforcement learning (IRL) deals with estimating an agent's utility function from its actions. In this paper, we consider how an agent can hide its strategy and mitigate…

cs.LG2021

Rationally Inattentive Utility Maximization for Interpretable Deep Image Classification

Kunal Pattanayak, Vikram Krishnamurthy

Are deep convolutional neural networks (CNNs) for image classification explainable by utility maximization with information acquisition costs? We demonstrate that deep CNNs behave…

cs.LG2020

Adaptive Non-reversible Stochastic Gradient Langevin Dynamics

Vikram Krishnamurthy, George Yin

It is well known that adding any skew symmetric matrix to the gradient of Langevin dynamics algorithm results in a non-reversible diffusion with improved convergence rate. This pap…

cs.LG20201 cited

A Markov Decision Process Approach to Active Meta Learning

Bingjia Wang, Alec Koppel, Vikram Krishnamurthy

In supervised learning, we fit a single statistical model to a given data set, assuming that the data is associated with a singular task, which yields well-tuned models for specifi…

cs.LG2020

Multi-kernel Passive Stochastic Gradient Algorithms and Transfer Learning

Vikram Krishnamurthy, George Yin

This paper develops a novel passive stochastic gradient algorithm. In passive stochastic approximation, the stochastic gradient algorithm does not have control over the location wh…

cs.LG2020

Langevin Dynamics for Adaptive Inverse Reinforcement Learning of Stochastic Gradient Algorithms

Vikram Krishnamurthy, George Yin

Inverse reinforcement learning (IRL) aims to estimate the reward function of optimizing agents by observing their response (estimates or actions). This paper considers IRL when noi…