collaborators

5 papers

stat.ML2026

Learning Kernel-Based MDPs from Episodic Preferential Feedback

Nikola Pavlovic, Sattar Vakili, Qing Zhao

Human feedback often arrives as preferences rather than calibrated numeric rewards, motivating reinforcement learning from preferential feedback, also referred to as reinforcement…

cs.LG2026

PeakFocus: Bridging Peak Localization and Intensity Regression via a Unified Multi-Scale Framework for Electricity Load Forecasting

Wangzhi Yu, Peng Zhu, Qing Zhao +2

Electricity load peak forecasting (ELPF), simultaneously predicting peak timing and intensity, is a prerequisite for effective grid scheduling and risk management. However, existin…

stat.ML2025

Differential Privacy in Kernelized Contextual Bandits via Random Projections

Nikola Pavlovic, Sudeep Salgia, Qing Zhao

We consider the problem of contextual kernel bandits with stochastic contexts, where the underlying reward function belongs to a known Reproducing Kernel Hilbert Space. We study th…

stat.ML2025

Differentially Private Kernelized Contextual Bandits

Nikola Pavlovic, Sudeep Salgia, Qing Zhao

We consider the problem of contextual kernel bandits with stochastic contexts, where the underlying reward function belongs to a known Reproducing Kernel Hilbert Space (RKHS). We s…

cs.LG2025

Characterizing the Accuracy-Communication-Privacy Trade-off in Distributed Stochastic Convex Optimization

Sudeep Salgia, Nikola Pavlovic, Yuejie Chi +1

We consider the problem of differentially private stochastic convex optimization (DP-SCO) in a distributed setting with clients, where each of them has a local dataset of i…