4 citations · 4 across the 5 of their papers we have counts for
13 papers
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
Kyuhee Kim, Auguste Poiroux, Antoine Bosselut
Formal verification guarantees proof validity but not formalization faithfulness. For natural-language logical reasoning, where models construct axiom systems from scratch without…
Large Language Models Align with the Human Brain during Creative Thinking
Mete Ismayilzada, Simone A. Luchini, Abdulkadir Gokce +5
Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely regarded as its core generative engin…
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
Mete Ismayilzada, Renqing Cuomao, Daniil Yurshevich +3
Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and commonsense knowledge, to discover insi…
AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI
Jérémy Barghorn, Anna Sotnikova, Sacha Friedli +1
Large-enrollment university courses face persistent challenges in providing timely and scalable instructional support. While generative AI holds promise, its effective use depends…
Evaluating Language Model Agency through Negotiations
Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski +4
We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings o…
Reliable Evaluation and Benchmarks for Statement Autoformalization
Auguste Poiroux, Gail Weiss, Viktor KunÄak +1
Evaluating statement autoformalization, translating natural language mathematics into formal languages like Lean 4, remains a significant challenge, with few metrics, datasets, and…