2 papers
cs.LG2026
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
Hong Xie, Xiao Hu, Tao Tan +5
The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also…
cs.AI2025
Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing
Hong Xie, Haoran Gu, Yanying Huang +2
This paper proposes a variant of multiple-play stochastic bandits tailored to resource allocation problems arising from LLM applications, edge intelligence, etc. The model is compo…