1 paper · 1 filter
Andrey Podivilov, Vadim Lomshakov, Sergey Savin +4
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people w…