2 papers
cs.AI2025
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
Yitong Cui, Liu Liu, Baosheng Yu +5
Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that a…
cs.SE2025
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
Yi Cui
We introduce WebApp1K, a novel benchmark for evaluating large language models (LLMs) in test-driven development (TDD) tasks, where test cases serve as both prompt and verification…