1 paper · 1 filter
Michael Hardy, Yunsung Kim
LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on benchmarks, downstream tasks, and, im…