activity
20242026
most citedMedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

10 citations · 10 across the 6 of their papers we have counts for

collaborators

11 papers

cs.LG2026

: Stratified Scaling Search for Test-Time in Diffusion Language Models

Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +6

Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without additional training. However, n…

cs.CY2026

Clinical Note Bloat Reduction for Efficient LLM Use

Jordan L. Cahoon, Chloe Stanwyck, Asad Aali +5

Health systems are rapidly deploying large language models (LLMs) that use clinical notes for clinical decision support applications. However, modern documentation practices rely h…

cs.LG2026

Attention Head Entropy of LLMs Predicts Answer Correctness

Sophie Ostmeier, Brian Axelrod, Maya Varma +6

Large language models (LLMs) often generate plausible yet incorrect answers, posing risks in safety-critical settings such as medicine. Human evaluation is expensive, and LLM-as-ju…

cs.CV2025

Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks

Arnav Singhvi, Vasiliki Bikia, Asad Aali +2

Vision-language foundation models (VLMs) show promise for diverse imaging tasks but often underperform on medical benchmarks. Prior efforts to improve performance include model fin…

cs.CL2025

Structured Prompts Improve Evaluation of Language Models

Asad Aali, Muhammad Ahmed Mohsin, Vasiliki Bikia +15

As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, framewo…

cs.CL2025

MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries

François Grolleau, Emily Alsentzer, Timothy Keyes +17

Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assur…