1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2025★ 1 cited
Medical Large Language Model Benchmarks Should Prioritize Construct Validity
Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini +4
Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation…
cs.AI2023
Estimating Uncertainty in Multimodal Foundation Models using Public Internet Data
Shiladitya Dutta, Hongbo Wei, Lars van der Laan +1
Foundation models are trained on vast amounts of data at scale using self-supervised learning, enabling adaptation to a wide range of downstream tasks. At test time, these models e…