1 paper · 1 filter
Chenyang Zhu, Spencer Hong, Jingyu Wu +6
The advent of complex, interconnected long-horizon LLM systems has made it incredibly tricky to identify where and when these systems break down. Evaluation capabilities that curre…