Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
Learning Kernel-Based MDPs from Episodic Preferential Feedback
Nikola Pavlovic, Sattar Vakili, Qing Zhao
Human feedback often arrives as preferences rather than calibrated numeric rewards, motivating reinforcement learning from preferential feedback, also referred to as reinforcement…
stat.ML2025
Differential Privacy in Kernelized Contextual Bandits via Random Projections
Nikola Pavlovic, Sudeep Salgia, Qing Zhao
We consider the problem of contextual kernel bandits with stochastic contexts, where the underlying reward function belongs to a known Reproducing Kernel Hilbert Space. We study th…
stat.ML2025
Differentially Private Kernelized Contextual Bandits
Nikola Pavlovic, Sudeep Salgia, Qing Zhao
We consider the problem of contextual kernel bandits with stochastic contexts, where the underlying reward function belongs to a known Reproducing Kernel Hilbert Space (RKHS). We s…