2 papers
cs.AI2026
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Alyssa Unell, Natalie Dullerud, Naomi Boneh +4
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignme…
cs.CE2025
STARC-9: A Large-scale Dataset for Multi-Class Tissue Classification for CRC Histopathology
Barathi Subramanian, Rathinaraja Jeyaraj, Mitchell Nevin Peterson +5
Multi-class tissue-type classification of colorectal cancer (CRC) histopathologic images is a significant step in the development of downstream machine learning models for diagnosi…