2 papers
cs.CV2026
MI-CXR: A Benchmark for Longitudinal Reasoning over Multi-Interval Chest X-rays
Sunghwan Steve Cho, Yunseok Han, Jaeyoung Do
Longitudinal chest X-ray (CXR) interpretation requires reasoning over disease evolution across multiple patient visits, yet most existing medical VQA benchmarks focus on single ima…
cs.AI2026
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
Yunseok Han, Yejoon Lee, Jaeyoung Do
Large Reasoning Models (LRMs) exhibit strong performance, yet often produce rationales that sound plausible but fail to reflect their true decision process, undermining reliability…