4 papers
Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning
Vincent Taboga, Justin Veilleux, Doseok Jang +2
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a criti…
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
Aneri Muni, Vincent Taboga, Esther Derman +2
Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events. Unlike risk-neutral object…
Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit Conservatism
Tianwei Ni, Esther Derman, Vineet Jain +3
Popular offline reinforcement learning (RL) methods rely on explicit conservatism, penalizing out-of-dataset actions or restricting rollout horizons. We question the universality o…
Discovery of Sustainable Refrigerants through Physics-Informed RL Fine-Tuning of Sequence Models
Adrien Goldszal, Diego Calanzone, Vincent Taboga +1
Most refrigerants currently used in air-conditioning systems, such as hydrofluorocarbons, are potent greenhouse gases and are being phased down. Large-scale molecular screening has…