◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Evan Frick

4 papers hereh-index 4570 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators

4 papers

cs.AI2026

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Yuanhao Ban, Tong Xie, Sohyun An +6

Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) models. Existing faithfulness ben…

cs.LG2025

Prompt-to-Leaderboard

Evan Frick, Connor Chen, Joseph Tennyson +4

Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and…

cs.LG2024

How to Evaluate Reward Models for RLHF

Evan Frick, Tianle Li, Connor Chen +6

We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-s…

cs.LG2024

From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Tianle Li, Wei-Lin Chiang, Evan Frick +5

The rapid evolution of Large Language Models (LLMs) has outpaced the development of model evaluation, highlighting the need for continuous curation of new, challenging benchmarks.…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.