◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Susu Zhang

3 papers hereh-index 217 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI2
  • stat.AP1

identity via Semantic Scholar / OpenAlex

works on
benchmark evaluation 1item response theory 1large language models 1model ranking 1simulation study 1

From the 1 of 3 linked papers with an AI index.

most citedAI Evaluation Should Require Standardized Item-Level Data Releases

1 citations · 1 across the 2 of their papers we have counts for

collaborators

3 papers

cs.AI2026

Can We Trust Item Response Theory for AI Evaluation?

Han Jiang, Sunbeom Kwon, Jinwen Luo +2

The paper investigates how well item response theory (IRT) works for evaluating large language model benchmarks, highlighting challenges when benchmark data differ from traditional…

cs.AI2026★ 1 cited

AI Evaluation Should Require Standardized Item-Level Data Releases

Han Jiang, Susu Zhang, Dongyao Zhu +6

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified it…

stat.AP2025

Reducing Differential Item Functioning via Process Data

Ling Chen, Susu Zhang, Jingchen Liu

Testing fairness is a major concern in psychometric and educational research. A typical approach for ensuring testing fairness is through differential item functioning (DIF) analys…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.