2 papers
cs.CL2026
MedVAL: Toward Expert-Level Medical Text Validation with Language Models
Asad Aali, Vasiliki Bikia, Maya Varma +24
With the growing use of language models (LMs) in clinical environments, there is an immediate need to evaluate the accuracy and safety of LM-generated medical text. Currently, such…
cs.LG2025
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
Olawale Salaudeen, Nicole Chiou, Shiny Weng +1
Spurious correlations, unstable statistical shortcuts a model can exploit, are expected to degrade performance out-of-distribution (OOD). However, across many popular OOD generaliz…