1 paper · 1 filter
Wenhao Li, Wenwu Li, Chuyun Shen +8
We present TextAtari, a benchmark for evaluating language agents on very long-horizon decision-making tasks spanning up to 100,000 steps. By translating the visual state representa…