#supervised fine-tuning
8 papers · 1 filter
Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
Site Li, Jianyi Hao, Xiaofeng Liu
The paper evaluates different adaptation strategies for multi-frame medical visual question answering and finds that a simple objective-aligned direct answer supervised fine-tuning…
Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
Anton de la Fuente, Arthur Conmy
The paper investigates whether lessons learned from supervised fine-tuning (SFT) in alignment training, model organisms, and toy models can be transferred across these domains, dem…
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit +2
The paper investigates why reinforcement‑learning‑trained models outperform supervised fine‑tuned models on math reasoning by analyzing their internal representations with linear p…
Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation
Yu-Du Feng, Niels Mündler-Sasahara, Mark Vero +1
The paper proposes a cheap method to adapt reasoning language models to new tasks by first instruction‑tuning them on ordinary supervised data and then merging the tuned model back…
Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion
Fengzhuo Zhang, Zhuoran Yang, Dirk Bergemann
The paper studies when users of large language models should use expensive supervised fine-tuning versus lightweight in‑context learning, considering how other users' choices creat…
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
Wangjin Zhou, Yizhou Zhang, Yichi Wang +1
The paper investigates how supervised fine-tuning performance for speech foundation models varies across different pretrained checkpoints, showing that gains often depend on the sp…