18 citations · 20 across the 2 of their papers we have counts for
7 papers
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
PersonaMail: Learning and Adapting Personal Communication Preferences for Context-Aware Email Writing
Rui Yao, Qiuyuan Ren, Felicia Fang-Yi Tan +3
LLM-assisted writing has seen rapid adoption in interpersonal communication, yet current systems often fail to capture the subtle tones essential for effectiveness. Email writing e…
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Elliot Glazer, Ege Erdil, Tamay Besiroglu +21
We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most…
Intrinsic Khovanov homology in
Qiuyu Ren, Hongjian Yang
We prove that Khovanov homology is an invariant of links in unparametrized 's, i.e., oriented -manifolds diffeomorphic to . Along the way, we estab…
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
Yue Zhang, Jiaxin Zhang, Qiuyu Ren +5
We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathe…
Cosmetic surgery on satellite knots
Qiuyu Ren
We show that if there exists a knot in that admits purely cosmetic surgeries, then there exists a hyperbolic one with this property.