activity
20242026
collaborators
Showing cs.LGShow all

11 papers · 1 filter

cs.LG2026

ICR-RL: Deep Reinforcement Learning via In-Context Regression

David Schiff, Ofir Lindenbaum, Yonathan Efroni

Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling them to generalize effectively to new, rela…

cs.LG2026

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

Amit Roth, Ankur Samanta, Matan Halevy +2

Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful und…

cs.LG2025

Imbalanced Gradients in RL Post-Training of Multi-Task LLMs

Runzhe Wu, Ankur Samanta, Ayush Jain +7

Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assume…

cs.LG2025

Simple Optimizers for Convex Aligned Multi-Objective Optimization

Ben Kretzu, Karen Ullrich, Yonathan Efroni

It is widely recognized in modern machine learning practice that access to a diverse set of tasks can enhance performance across those tasks. This observation suggests that, unlike…

cs.LG2025

Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do

Yoav Wald, Mark Goldstein, Yonathan Efroni +2

Problems in fields such as healthcare, robotics, and finance requires reasoning about the value both of what decision or action to take and when to take it. The prevailing hope is…

cs.LG2025

Aligned Multi Objective Optimization

Yonathan Efroni, Ben Kretzu, Daniel Jiang +4

To date, the multi-objective optimization literature has mainly focused on conflicting objectives, studying the Pareto front, or requiring users to balance tradeoffs. Yet, in machi…