1 paper
Rongge Xu, Hui Dai, Yiming Fu +5
While large language models (LLMs) have demonstrated impressive capabilities in formal theorem proving, current benchmarks fail to adequately measure library-grounded abstraction -…