4 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
Load Balancing Mixture of Experts with Similarity Preserving Routers
Nabil Omi, Siddhartha Sen, Ali Farhadi
Sparse Mixture of Experts (MoE) models offer a scalable and efficient architecture for training large neural networks by activating only a subset of parameters ("experts") for each…
Generative Modeling of Individual Behavior at Scale
Nabil Omi, Lucas Caccia, Anurag Sarkar +2
There has been a growing interest in using AI to model human behavior, particularly in domains where humans interact with this technology. While most existing work models human beh…
Progressive Safeguards for Safe and Model-Agnostic Reinforcement Learning
Nabil Omi, Hosein Hasanbeig, Hiteshi Sharma +2
In this paper we propose a formal, model-agnostic meta-learning framework for safe reinforcement learning. Our framework is inspired by how parents safeguard their children across…