From the 1 of 6 linked papers with an AI index.
6 papers
Can We Trust Item Response Theory for AI Evaluation?
Han Jiang, Sunbeom Kwon, Jinwen Luo +2
The paper investigates how well item response theory (IRT) works for evaluating large language model benchmarks, highlighting challenges when benchmark data differ from traditional…
AI Evaluation Should Require Standardized Item-Level Data Releases
Han Jiang, Susu Zhang, Dongyao Zhu +6
This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…
The Impact of Generative AI on Architectural Conceptual Design: Performance, Creative Self-Efficacy and Cognitive Load
Han Jiang, Yao Xiao, Rachel Hurley +1
Our study examines how generative AI (GenAI) influences performance, creative self-efficacy, and cognitive load in architectural conceptual design tasks. Thirty-six student partici…
Human-AI Narrative Synthesis to Foster Shared Understanding in Civic Decision-Making
Cassandra Overney, Hang Jiang, Urooj Haider +6
Community engagement processes in representative political contexts, like school districts, generate massive volumes of feedback that overwhelm traditional synthesis methods, creat…
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
Han Jiang, Pengda Wang, Xiaoyuan Yi +2
Social sciences have accumulated a rich body of theories and methodologies for investigating the human mind and behaviors, while offering valuable insights into the design and unde…
Automatic Detection of Research Values from Scientific Abstracts Across Computer Science Subfields
Hang Jiang, Tal August, Luca Soldaini +2
The field of Computer science (CS) has rapidly evolved over the past few decades, providing computational tools and methodologies to various fields and forming new interdisciplinar…