66 citations · 104 across the 10 of their papers we have counts for
5 papers · 1 filter
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
Runzhe Wu, Ankur Samanta, Ayush Jain +7
Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assume…
Simple Optimizers for Convex Aligned Multi-Objective Optimization
Ben Kretzu, Karen Ullrich, Yonathan Efroni
It is widely recognized in modern machine learning practice that access to a diverse set of tasks can enhance performance across those tasks. This observation suggests that, unlike…
Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
Ankur Samanta, Akshayaa Magesh, Runzhe Wu +7
Self-improvement, where models improve beyond their current performance without external supervision, remains a challenge. The core difficulty is sourcing a training signal stronge…
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
Yoav Wald, Mark Goldstein, Yonathan Efroni +2
Problems in fields such as healthcare, robotics, and finance requires reasoning about the value both of what decision or action to take and when to take it. The prevailing hope is…
Aligned Multi Objective Optimization
Yonathan Efroni, Ben Kretzu, Daniel Jiang +4
To date, the multi-objective optimization literature has mainly focused on conflicting objectives, studying the Pareto front, or requiring users to balance tradeoffs. Yet, in machi…