3 papers
cs.CV2026
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
Ashwin Kumar, Robbie Holland, Corey Barrett +8
Recent medical multimodal foundation models are built as multimodal LLMs (MLLMs) by connecting a CLIP-pretrained vision encoder to an LLM using LLaVA-style finetuning. This two-sta…
cs.MA2026
RadAgents: Multimodal Agentic Reasoning for Chest X-ray Interpretation with Radiologist-like Workflows
Kai Zhang, Corey D Barrett, Jangwon Kim +3
Agentic systems offer a potential path to solve complex clinical tasks through collaboration among specialized agents, augmented by tool use and external knowledge bases. Neverthel…
cs.CL2025
Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
Samyak Jhaveri, Praphul Singh, Jangwon Kim +2
Automating clinical documentation with large language models requires precise alignment with priorities such as completeness and factual grounding. We present an evaluation-integra…