8 papers
SciR: A Controllable Benchmark for Scientific Reasoning in LLMs
Pierre Beckmann, Marco Valentino, Andre Freitas
Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction. Reliably evaluating LLMs on these in scientific settings is cur…
AbstRAG: Learning to Abstract for Retrieval Problems
Lei Xu, Xin Quan, Daniel Pedronette +1
Retrieval-augmented generation often fails when the query, the document evidence, and the user's intent are expressed at different levels of abstraction. A query may ask about a cl…
Reasoning without Gold Standards: A Proxy-Judge Theory of Autoformalization
Lei Xu, Xin Quan, André Freitas
Complex reasoning tasks increasingly require systems to produce outputs whose correctness cannot be judged by exact match against a single reference. Autoformalization (AF) is a re…
Monotonic Reference-Free Refinement for Autoformalization
Lan Zhang, Marco Valentino, André Freitas
While statement autoformalization has advanced rapidly, full-theorem autoformalization remains largely unexplored. Existing iterative refinement methods in statement autoformalizat…
Decompose-and-Formalise: Recursively Verifiable Natural Language Inference
Xin Quan, Marco Valentino, Louise A. Dennis +1
Recent work has shown that integrating large language models (LLMs) with theorem provers (TPs) in neuro-symbolic pipelines helps with entailment verification and proof-guided refin…
MASA: LLM-Driven Multi-Agent Systems for Autoformalization
Lan Zhang, Marco Valentino, André Freitas
Autoformalization serves a crucial role in connecting natural language and formal reasoning. This paper presents MASA, a novel framework for building multi-agent systems for autofo…