1 paper · 1 filter
Miguel Zabaleta, Baihan Lin
Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolati…