5 papers
WebDesignIter: Co-Evolving Design Knowledge for Repository-Level Front-End Code Generation
Zheng Pei, Mingwei Liu, Zhenxi Chen +2
Front-end development accumulates change after change at the repository level, weaving complex cross-file dependencies that current LLM coding agents tuned for single-shot tasks ca…
Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition
Mingwei Liu, Zhenxi Chen, Zheng Pei +3
Implementing new features across an entire codebase presents a formidable challenge for Large Language Models (LLMs). This proactive task requires a deep understanding of the globa…
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
Yimin Zhao, Sheela R. Damle, Simone E. Dekker +13
Large language models (LLMs) have achieved expert-level performance on standardized examinations, yet multiple-choice accuracy poorly reflects real-world clinical utility and safet…
Humanity's Last Exam
Long Phan, Alice Gatti, Ziwen Han +1144
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…
UQ: Assessing Language Models on Unsolved Questions
Fan Nie, Ken Ziyu Liu, Zihao Wang +11
Benchmarks shape progress in AI research. A useful benchmark should be both difficult and realistic: questions should challenge frontier models while also reflecting real-world usa…