activity
20182025
most citedAction Robust Reinforcement Learning and Applications in Continuous Control

66 citations · 104 across the 10 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.LG2025

Imbalanced Gradients in RL Post-Training of Multi-Task LLMs

Runzhe Wu, Ankur Samanta, Ayush Jain +7

Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assume…

cs.LG2025

Simple Optimizers for Convex Aligned Multi-Objective Optimization

Ben Kretzu, Karen Ullrich, Yonathan Efroni

It is widely recognized in modern machine learning practice that access to a diverse set of tasks can enhance performance across those tasks. This observation suggests that, unlike…

cs.AI2025

Self-Improvement of Language Models by Post-Training on Multi-Agent Debate

Ankur Samanta, Akshayaa Magesh, Runzhe Wu +7

Self-improvement, where models improve beyond their current performance without external supervision, remains a challenge. The core difficulty is sourcing a training signal stronge…

cs.LG2025

Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do

Yoav Wald, Mark Goldstein, Yonathan Efroni +2

Problems in fields such as healthcare, robotics, and finance requires reasoning about the value both of what decision or action to take and when to take it. The prevailing hope is…

cs.LG2025

Aligned Multi Objective Optimization

Yonathan Efroni, Ben Kretzu, Daniel Jiang +4

To date, the multi-objective optimization literature has mainly focused on conflicting objectives, studying the Pareto front, or requiring users to balance tradeoffs. Yet, in machi…