1 paper · 1 filter
QuocViet Pham, Elvir Karimov, Andrey Galichin +1
LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style problems and often fail to c…