1 paper
Yilong Li, Suman Banerjee, Tong Che
Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@K as the canonical metric. Yet the standard policy class draws…