3 papers
cs.CL2025
Search Arena: Analyzing Search-Augmented LLMs
Mihran Miroyan, Tsung-Han Wu, Logan King +8
Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains chall…
cs.LG2025
Prompt-to-Leaderboard
Evan Frick, Connor Chen, Joseph Tennyson +4
Large language model (LLM) evaluations typically rely on aggregated metrics like accuracy or human preference, averaging across users and prompts. This averaging obscures user- and…
cs.LG2024
How to Evaluate Reward Models for RLHF
Evan Frick, Tianle Li, Connor Chen +6
We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-s…