activity
20212026
most citedScalable Distributional Robustness in a Class of Non Convex Optimization with Guarantees

1 citations · 1 across the 8 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2025

LoRe: Personalizing LLMs via Low-Rank Reward Modeling

Avinandan Bose, Zhihan Xiong, Yuejie Chi +3

Personalizing large language models (LLMs) to accommodate diverse user preferences is essential for enhancing alignment and user satisfaction. Traditional reinforcement learning fr…

cs.LG2025

Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning

Avinandan Bose, Laurent Lessard, Maryam Fazel +1

The rise of foundation models fine-tuned on human feedback from potentially untrusted users has increased the risk of adversarial data poisoning, necessitating the study of robustn…

cs.LG2024

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration

Avinandan Bose, Zhihan Xiong, Aadirupa Saha +2

Reinforcement Learning from Human Feedback (RLHF) is currently the leading approach for aligning large language models with human preferences. Typically, these models rely on exten…

cs.LG2024

Offline Multi-task Transfer RL with Representational Penalization

Avinandan Bose, Simon Shaolei Du, Maryam Fazel

We study the problem of representation transfer in offline Reinforcement Learning (RL), where a learner has access to episodic data from a number of source tasks collected a priori…

cs.LG2023

Initializing Services in Interactive ML Systems for Diverse Users

Avinandan Bose, Mihaela Curmei, Daniel L. Jiang +4

This paper investigates ML systems serving a group of users, with multiple models/services, each aimed at specializing to a sub-group of users. We consider settings where upon depl…

cs.LG20221 cited

Scalable Distributional Robustness in a Class of Non Convex Optimization with Guarantees

Avinandan Bose, Arunesh Sinha, Tien Mai

Distributionally robust optimization (DRO) has shown lot of promise in providing robustness in learning as well as sample based optimization problems. We endeavor to provide DRO so…