4 papers
Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment
Xun Shen, Yuepeng Wang, Akifumi Wachi +13
Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical inte…
MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning
Yuepeng Wang, Ken Kawano, Yongqi Zhou +11
Medical treatment recommendation poses several challenges to reinforcement learning (RL): patient physiology evolves in continuous time, measurements and interventions are performe…
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
Hong Xie, Xiao Hu, Tao Tan +5
The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also…
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
Xiao Hu, Hong Xie, Tao Tan +2
A large number of heuristics have been proposed to optimize the reinforcement fine-tuning of LLMs. However, inconsistent claims are made from time to time, making this area elusive…