1 paper
QuocViet Pham, Elvir Karimov, Andrey Galichin +1
LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to c…