2 papers
cs.LG2026
SCALAR: Learning and Composing Skills through LLM Guided Symbolic Planning and Deep RL Grounding
Renos Zabounidis, Yue Wu, Simon Stepputtis +4
LM-based agents excel when given high-level action APIs but struggle to ground language into low-level control. Prior work has LLMs generate skills or reward functions for RL, but…
stat.ML2026
POLAR: A Pessimistic Model-based Policy Learning Algorithm for Dynamic Treatment Regimes
Ruijia Zhang, Xiangyu Zhang, Zhengling Qi +2
Dynamic treatment regimes (DTRs) provide a principled framework for optimizing sequential decision-making in domains where decisions must adapt over time in response to individual…