2 papers
cs.LG2026
Efficient Opportunistic Approachability
Teodor Vanislavov Marinov, Mehryar Mohri, Princewill Okoroafor +2
We study the problem of opportunistic approachability: a generalization of Blackwell approachability where the learner would like to obtain stronger guarantees (i.e., approach a sm…
cs.LG2025
Design Considerations in Offline Preference-based RL
Alekh Agarwal, Christoph Dann, Teodor V. Marinov
Offline algorithms for Reinforcement Learning from Human Preferences (RLHF), which use only a fixed dataset of sampled responses given an input, and preference feedback among these…