4 papers
Vector-Bench: Can Models Surgically Edit SVG Code?
Yug Aditi Gupta, Prannay Hebbar
Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone. The second is easy to miss when an output is judged only as…
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai, Jesse Hu, Joan Cabezas +23
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent b…
SIA: Self Improving AI with Harness & Weight Updates
Prannay Hebbar, Yogendra Manawat, Samuel Verboomen +4
Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The long-horizon goal of an AI th…
REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real Websites
Divyansh Garg, Shaun VanWeelden, Diego Caples +15
We introduce REAL, a benchmark and framework for multi-turn agent evaluations on deterministic simulations of real-world websites. REAL comprises high-fidelity, deterministic repli…