2 papers
cs.CL2026
RadOT-Eval: Auditable Structured-Evidence Transport for Radiology Report Evaluation
Weixin Liu, Juming Xiong, Yang Li +5
Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertain…
cs.LG2026
RiskNet: A large-scale dataset of AI risk incidents from news with alignment and multi-dimensional annotations
Leihan Zhang, Wecheng Ye, Xianlong Ma +5
As artificial intelligence (AI) systems are increasingly deployed across socially consequential domains, reports of AI-related harms and failures have grown in frequency and divers…