Publications (4)
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
Application-Level Validation of Accelerator Designs Using a Formal Software/Hardware Interface
Bo-Yuan Huang, Steven Lyubomirsky, Yi Li +10
Ideally, accelerator development should be as easy as software development. Several recent design languages/tools are working toward this goal, but actually testing early designs o…
Dynamic Tensor Rematerialization
Marisa Kirisame, Steven Lyubomirsky, Altan Haan +5
Checkpointing enables the training of deep learning models under restricted memory budgets by freeing intermediate activations from memory and recomputing them on demand. Current c…
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
Marcus J. Min, Mike He, Zhaoyu Li +5
The paper proposes shifting autoformalization from isolated statements to theory-level, aiming to automatically translate whole bodies of mathematical knowledge—including axioms, d…