2 papers
cs.LG2026
Smooth Multi-Policy Causal Effect Estimation in Longitudinal Settings
Wenxin Chen, Weishen Pan, Kyra Gan +1
Comparative evaluation of multiple dynamic treatment policies is essential for healthcare and policy decisions, yet conventional longitudinal causal inference methods estimate each…
cs.LG2026
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
Nirmal Patel, Fei Wang, Inderjit S. Dhillon
The alignment of Large Language Models (LLMs) utilizes Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form question answering and open-ende…