1 citations · 1 across the 4 of their papers we have counts for
Showing cs.SEShow all
3 papers · 1 filter
cs.SE2025★ 1 cited
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Yi Cui
We introduce WebApp1K, a novel benchmark for evaluating large language models (LLMs) in test-driven development (TDD) tasks, where test cases serve as both prompt and verification…
cs.SE2024
A Case Study of Web App Coding with OpenAI Reasoning Models
Yi Cui
This paper presents a case study of coding tasks by the latest reasoning models of OpenAI, i.e. o1-preview and o1-mini, in comparison with other frontier models. The o1 models deli…
cs.SE2024
Insights from Benchmarking Frontier Language Models on Web App Code Generation
Yi Cui
This paper presents insights from evaluating 16 frontier large language models (LLMs) on the WebApp1K benchmark, a test suite designed to assess the ability of LLMs to generate web…