human feedback 1model exploitation 1offline reinforcement learning 1preference learning 1world models 1
From the 1 of 25 linked papers with an AI index.
3 citations · 3 across the 13 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
The paper introduces RENEW, a method that uses human preferences over imagined rollouts to correct model exploitation in offline model-based reinforcement learning, focusing fine‑t…
cs.LG2025
Beyond Discriminant Patterns: On the Robustness of Decision Rule Ensembles
Xin Du, Subramanian Ramamoorthy, Wouter Duivesteijn +2
Local decision rules are commonly understood to be more explainable, due to the local nature of the patterns involved. With numerical optimization methods such as gradient boosting…