1 paper
Chuhan Li, Ziyao Shangguan, Yilun Zhao +3
Existing benchmarks for evaluating foundation models mainly focus on single-document, text-only tasks. However, they often fail to fully capture the complexity of research workflow…