From the 1 of 7 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MIRA: A Bilingual Benchmark for Medical Information Response Audit
Mengyu Xu, Qiaoxin Yang, Qianqian Wang +3
Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medic…
cs.AI2026
Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
Chongyang Gao, Diji Yang, Shuyan Zhou +4
We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench…