4 papers
Model Specific Task Similarity for Vision Language Model Selection via Layer Conductance
Wei Yang, Hong Xie, Tao Tan +3
While open sourced Vision-Language Models (VLMs) have proliferated, selecting the optimal pretrained model for a specific downstream task remains challenging. Exhaustive evaluation…
Demystifying Design Choices of Reinforcement Fine-tuning: A Batched Contextual Bandit Learning Perspective
Hong Xie, Xiao Hu, Tao Tan +5
The reinforcement fine-tuning area is undergoing an explosion papers largely on optimizing design choices. Though performance gains are often claimed, inconsistent conclusions also…
Rethinking Reinforcement fine-tuning of LLMs: A Multi-armed Bandit Learning Perspective
Xiao Hu, Hong Xie, Tao Tan +2
A large number of heuristics have been proposed to optimize the reinforcement fine-tuning of LLMs. However, inconsistent claims are made from time to time, making this area elusive…
Multiple-play Stochastic Bandits with Prioritized Arm Capacity Sharing
Hong Xie, Haoran Gu, Yanying Huang +2
This paper proposes a variant of multiple-play stochastic bandits tailored to resource allocation problems arising from LLM applications, edge intelligence, etc. The model is compo…