2 papers
cs.LG2026
torchtune: PyTorch native post-training library
Mark Obozov, Maxime Griot, Joseph Cummings +8
Modern LLMs typically require multistage training pipelines to achieve strong downstream performance, with post-training serving as the main interface for adapting open-weight mode…
cs.CL2025
Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis
Felipe Ribeiro Fujita de Mello, Hideyuki Takada
We investigated the impact of data selection on machine translation fine-tuning for open LLMs. Using Japanese-English corpora, we compare five selectors: TF-IDF, COMET Kiwi, QuRate…