From the 1 of 7 linked papers with an AI index.
7 papers
Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO
Mengyu Xu, Qiaoxin Yang, Zhihan Liu +4
Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different q…
Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
Shuyi Fan, Boyuan Deng, Mengyu Xu +4
The paper audits whether a general helpfulness rubric can reliably distinguish between answer‑giving and pedagogical guidance in large language model tutoring, finding that helpful…
MIRA: A Bilingual Benchmark for Medical Information Response Audit
Mengyu Xu, Qiaoxin Yang, Qianqian Wang +3
Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether responses preserve comparable medic…
Learning Spatial-Preserving Hierarchical Representations for Digital Pathology
Weiyi Wu, Xingjian Diao, Chunhui Zhang +4
Whole slide images (WSIs) pose fundamental computational challenges due to their gigapixel resolution and the sparse distribution of informative regions. Existing approaches often…
Classroom Final Exam: An Instructor-Tested Reasoning Benchmark
Chongyang Gao, Diji Yang, Shuyan Zhou +4
We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench…
Exploiting Label-Independent Regularization from Spatial Dependencies for Whole Slide Image Analysis
Weiyi Wu, Xinwen Xu, Chongyang Gao +3
Whole slide images, with their gigapixel-scale panoramas of tissue samples, are pivotal for precise disease diagnosis. However, their analysis is hindered by immense data size and…