1 paper · 1 filter
Jongyoon Kim, Hojae Han, Seung-won Hwang
Recent advances in large language models (LLMs) have shown promise in formal theorem proving, yet evaluating semantic correctness remains challenging. Existing evaluations rely on…