9 citations · 9 across the 10 of their papers we have counts for
4 papers · 1 filter
SHADOWBENCH: Toward Reliable Automatic Evaluation of Semantic Alignment in Autoformalization
Hojae Han, Jongyoon Kim, Sanghyeok Park +8
Autoformalization translates informal mathematical theorems into code for proof assistants such as Lean. A central challenge is that current evaluation metrics can accept type-corr…
Benchmarking Testing in Automated Theorem Proving
Jongyoon Kim, Hojae Han, Seung-won Hwang
Recent advances in large language models (LLMs) have shown promise in formal theorem proving, yet evaluating semantic correctness remains challenging. Existing evaluations rely on…
Can David Beat Goliath? On Multi-Hop Reasoning with Resource-Constrained Agents
Hojae Han, Heeyun Jung, Jongyoon Kim +1
Multi-turn reasoning agents solve complex questions by decomposing them into intermediate retrieval or tool-use steps, for accumulating supporting evidence across turns. Meanwhile,…
Arctic-SnowCoder: Demystifying High-Quality Data in Code Pretraining
Yuxiang Wei, Hojae Han, Rajhans Samdani
Recent studies have been increasingly demonstrating that high-quality data is crucial for effective pretraining of language models. However, the precise definition of "high-quality…