2 papers
cs.CV2026
NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding
Mohammad H. Abbasi, Favour Nerrise, Shaurnav Ghosh +12
We present NeuroQA, a large-scale benchmark for visual question answering in 3D brain magnetic resonance imaging (MRI), with 56,953 QA pairs from 12,977 subjects across 12 datasets…
cs.CL2026
TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots
Fangrui Huang, Souhad Chbeir, Arpandeep Khatua +8
Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks-…