Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Endless Terminals: Scaling RL Environments for Terminal Agents
Kanishk Gandhi, Shivam Garg, Noah D. Goodman +1
Environments are the bottleneck for self-improving agents. Current terminal benchmarks were built for evaluation, not training; reinforcement learning requires a scalable pipeline,…
cs.LG2025
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
Kanishk Gandhi, Michael Y. Li, Lyle Goodyear +5
Understanding the world and explaining it with scientific theories is a central aspiration of artificial intelligence research. Proposing theories, designing experiments to test th…
cs.LG2024
CriticAL: Critic Automation with Language Models
Michael Y. Li, Vivek Vajipey, Noah D. Goodman +1
Understanding the world through models is a fundamental goal of scientific research. While large language model (LLM) based approaches show promise in automating scientific discove…