2 papers
cs.CL2025
Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference Data
Siqi Guo, Ilgee Hong, Vicente Balmaseda +6
Supervised fine-tuning (SFT) has become a crucial step for aligning pretrained large language models (LLMs) using supervised datasets of input-output pairs. However, despite being…
cs.LG2024
Stochastic Approximation Approaches to Group Distributionally Robust Optimization and Beyond
Lijun Zhang, Haomin Bai, Peng Zhao +2
This paper investigates group distributionally robust optimization (GDRO) with the goal of learning a model that performs well over different distributions. First, we formulate…