3 papers
cs.AI2026
AI Evaluation Should Work With Humans
Jan Kulveit, Gavin Leech, Tomáš Gavenčiak +1
This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) i…
physics.med-ph2026
Feasibility of a General-Purpose Deep Learning Dose Engine: A Multi-Site Validation Study
Yao Zhao, Ka Ho Tam, Raphael Douglas +6
Conventional radiotherapy dose calculation algorithms are often computationally slow and non-differentiable, creating bottlenecks for online adaptive radiotherapy (ART) and limitin…
cs.CV2024
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
Harshit, Tolga Tasdizen
The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Mode…