1 citations · 1 across the 4 of their papers we have counts for
4 papers
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
Yitong Cui, Liu Liu, Baosheng Yu +5
Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that a…
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Yi Cui
We introduce WebApp1K, a novel benchmark for evaluating large language models (LLMs) in test-driven development (TDD) tasks, where test cases serve as both prompt and verification…
A Case Study of Web App Coding with OpenAI Reasoning Models
Yi Cui
This paper presents a case study of coding tasks by the latest reasoning models of OpenAI, i.e. o1-preview and o1-mini, in comparison with other frontier models. The o1 models deli…
Insights from Benchmarking Frontier Language Models on Web App Code Generation
Yi Cui
This paper presents insights from evaluating 16 frontier large language models (LLMs) on the WebApp1K benchmark, a test suite designed to assess the ability of LLMs to generate web…