2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Kaivalya Hariharan, Uzay Girit, Atticus Wang +1
Benchmarks for large language models (LLMs) have predominantly assessed short-horizon, localized reasoning. Existing long-horizon suites (e.g. SWE-bench) rely on manually curated i…