most citedHumanity's Last Exam

18 citations · 20 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG202618 cited

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.HC20262 cited

PersonaMail: Learning and Adapting Personal Communication Preferences for Context-Aware Email Writing

Rui Yao, Qiuyuan Ren, Felicia Fang-Yi Tan +3

LLM-assisted writing has seen rapid adoption in interpersonal communication, yet current systems often fail to capture the subtle tones essential for effectiveness. Email writing e…

cs.AI2025

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Elliot Glazer, Ege Erdil, Tamay Besiroglu +21

We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most…

math.GT2025

Intrinsic Khovanov homology in

Qiuyu Ren, Hongjian Yang

We prove that Khovanov homology is an invariant of links in unparametrized 's, i.e., oriented -manifolds diffeomorphic to . Along the way, we estab…

cs.AI2025

GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models

Yue Zhang, Jiaxin Zhang, Qiuyu Ren +5

We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathe…

math.GT2025

Cosmetic surgery on satellite knots

Qiuyu Ren

We show that if there exists a knot in that admits purely cosmetic surgeries, then there exists a hyperbolic one with this property.