2 papers
cs.LG2026
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
Hong Xie, Xiao Hu, Tao Tan +5
The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also…
cs.LG2026
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
Xiao Hu, Hong Xie, Tao Tan +2
A large number of heuristics have been proposed to optimize the reinforcement fine-tuning of LLMs. However, inconsistent claims are made from time to time, making this area elusive…