2 papers
cs.LG2026
VERPO: Verified Evidence Regularized Policy Optimization
Haijiang Li, Chengyu Lv, Yi Zhang +8
Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-co…
cs.CV2025
MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation
Junjie Yang, Yuhao Yan, Gang Wu +12
Medical vision-language models (VLMs) are increasingly expected to support clinical workflows through diagnostic text and relevant medical images. However, current medical visual b…