3 papers
cs.GT2026
Efficient Decentralized Learning of Generalized Quantal Response Equilibrium
Zehao Zhao, Apurv Shukla, Rahul Jain +1
We study a solution concept for bounded rational agents in finite normal-form general-sum games called Generalized Quantal Response Equilibrium (GQRE) which generalizes Quantal Res…
cs.LG2025
FraPPE: Fast and Efficient Preference-based Pure Exploration
Udvas Das, Apurv Shukla, Debabrota Basu
Preference-based Pure Exploration (PrePEx) aims to identify with a given confidence level the set of Pareto optimal arms in a vector-valued (aka multi-objective) bandit, where the…
cs.LG2025
Vector preference-based contextual bandits under distributional shifts
Apurv Shukla, P. R. Kumar
We consider contextual bandit learning under distribution shift when reward vectors are ordered according to a given preference cone. We propose an adaptive-discretization and opti…