3 papers
cs.CL2026
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
Ajmal M., Abin Roy, Afthab Salam Kanniyan +4
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…
cs.CL2026
BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
Mathew J. Koretsky, Maya Willey, Owen Bianchi +7
Biomedical researchers increasingly rely on large-scale structured databases for complex analytical tasks. However, current text-to-SQL systems often struggle to map qualitative sc…
cs.CL2025
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
Owen Bianchi, Mathew J. Koretsky, Maya Willey +7
Large language models (LLMs) face significant challenges with needle-in-ahaystack tasks, where relevant information ("the needle") must be drawn from a large pool of irrelevant con…