activity
20242026
most citedEvaluating Language Model Agency through Negotiations

4 citations · 4 across the 5 of their papers we have counts for

collaborators

13 papers

cs.AI2026

Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning

Kyuhee Kim, Auguste Poiroux, Antoine Bosselut

Formal verification guarantees proof validity but not formalization faithfulness. For natural-language logical reasoning, where models construct axiom systems from scratch without…

q-bio.NC2026

Large Language Models Align with the Human Brain during Creative Thinking

Mete Ismayilzada, Simone A. Luchini, Abdulkadir Gokce +5

Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is widely regarded as its core generative engin…

cs.CL2026

CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge

Mete Ismayilzada, Renqing Cuomao, Daniil Yurshevich +3

Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and commonsense knowledge, to discover insi…

cs.CY2026

AI Meets Mathematics Education: A Case Study on Supporting an Instructor in a Large Mathematics Class with Context-Aware AI

Jérémy Barghorn, Anna Sotnikova, Sacha Friedli +1

Large-enrollment university courses face persistent challenges in providing timely and scalable instructional support. While generative AI holds promise, its effective use depends…

cs.CL20264 cited

Evaluating Language Model Agency through Negotiations

Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski +4

We introduce an approach to evaluate language model (LM) agency using negotiation games. This approach better reflects real-world use cases and addresses some of the shortcomings o…

cs.CL2025

Reliable Evaluation and Benchmarks for Statement Autoformalization

Auguste Poiroux, Gail Weiss, Viktor Kunčak +1

Evaluating statement autoformalization, translating natural language mathematics into formal languages like Lean 4, remains a significant challenge, with few metrics, datasets, and…