Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Atria Dawn: The Dawn of Agentic Superintelligence
Honglin Guo, Tao Gui, Yicheng Chen +139
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn…
cs.AI2026
SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
Pujun Zheng, Zixin Shang, Shufan Jiang +5
SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluat…