4 papers · 1 filter
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Yi Cui
We introduce WebApp1K, a novel benchmark for evaluating large language models (LLMs) in test-driven development (TDD) tasks, where test cases serve as both prompt and verification…
A Case Study of Web App Coding with OpenAI Reasoning Models
Yi Cui
This paper presents a case study of coding tasks by the latest reasoning models of OpenAI, i.e. o1-preview and o1-mini, in comparison with other frontier models. The o1 models deli…
Insights from Benchmarking Frontier Language Models on Web App Code Generation
Yi Cui
This paper presents insights from evaluating 16 frontier large language models (LLMs) on the WebApp1K benchmark, a test suite designed to assess the ability of LLMs to generate web…
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
Yi Cui
We introduce WebApp1K, a practical code-generation benchmark to measure LLM ability to develop web apps. This benchmark aims to calibrate LLM output and aid the models to progressi…