◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Michael Hardy

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author2
  • middle author1

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.AI1

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.AI2025

Fantastic Bugs and Where to Find Them in AI Benchmarks

Sang Truong, Yuheng Tu, Michael Hardy +8

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of…

cs.CL2025

Measuring Teaching with LLMs

Michael Hardy

Objective and scalable measurement of teaching quality is a persistent challenge in education. While Large Language Models (LLMs) offer potential, general-purpose models have strug…

cs.CL2024

"All that Glitters": Approaches to Evaluations with Unreliable Model and Human Annotations

Michael Hardy

"Gold" and "ground truth" human-mediated labels have error. The effects of this error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias,…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.