From the 1 of 1 linked paper with an AI index.
1 paper
Chao Peng, Zhiheng Lyu, Peijie Dong +2
The paper proposes a benchmark metric called the horizon residual to compare long-horizon task success against predictions from short-stage baselines, highlighting how performance…