1 paper
Lech Madeyski, Barbara Kitchenham, Martin Shepperd
Context: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can m…