1 paper · 1 filter
Ruitong Li, Binjie Guo, Aisheng Mo +4
Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable software. Prompt-driven "vibe…