2 papers
cs.LG2025
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
Runzhe Wu, Ankur Samanta, Ayush Jain +7
Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assume…
cs.LG2025
Simple Optimizers for Convex Aligned Multi-Objective Optimization
Ben Kretzu, Karen Ullrich, Yonathan Efroni
It is widely recognized in modern machine learning practice that access to a diverse set of tasks can enhance performance across those tasks. This observation suggests that, unlike…