8 papers
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
Johannes Moll, Jean-Philippe Corbeil, Jiazhen Pan +4
LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-imp…
VERT: Reliable LLM Judges for Radiology Report Evaluation
Federica Bologna, Jean-Philippe Corbeil, Matthew Wilkens +1
Current literature on radiology report evaluation has focused primarily on designing LLM-based metrics and fine-tuning small models for chest X-rays. However, it remains unclear wh…
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
Ji Young Byun, Young-Jin Park, Jean-Philippe Corbeil +1
As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical…
Less Finetuning, Better Retrieval: Rethinking LLM Adaptation for Biomedical Retrievers via Synthetic Data and Model Merging
Sameh Khattab, Jean-Philippe Corbeil, Osman Alperen KoraÅ +5
Retrieval-augmented generation (RAG) has become the backbone of grounding Large Language Models (LLMs), improving knowledge updates and reducing hallucinations. Recently, LLM-based…
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot +4
As the performance of large language models (LLMs) continues to advance, their adoption in the medical domain is increasing. However, most existing risk evaluations largely focused…
Overview of the MEDIQA-OE 2025 Shared Task on Medical Order Extraction from Doctor-Patient Consultations
Jean-Philippe Corbeil, Asma Ben Abacha, Jerome Tremblay +4
Clinical documentation increasingly uses automatic speech recognition and summarization, yet converting conversations into actionable medical orders for Electronic Health Records r…