3 papers
cs.AI2026
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Jiwan Chung, JiHyuk Byun, Vibhav Vineet +1
Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improv…
cs.CV2026
Overconfidence and Calibration in Medical VQA: Empirical Findings and Hallucination-Aware Mitigation
Ji Young Byun, Young-Jin Park, Jean-Philippe Corbeil +1
As vision-language models (VLMs) are increasingly deployed in clinical decision support, more than accuracy is required: knowing when to trust their predictions is equally critical…
cs.CV2025
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
Ji Young Byun, Young-Jin Park, Navid Azizan +1
As a cornerstone of patient care, clinical decision-making significantly influences patient outcomes and can be enhanced by large language models (LLMs). Although LLMs have demonst…