5 papers
Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction
Serafim Batzoglou
First-order concept synthesis asks a system to infer one formula that classifies labeled objects consistently across several finite relational structures. Every candidate can be ev…
INDUCTION: Finite-Structure Concept Synthesis in First-Order Logic
Serafim Batzoglou
We introduce INDUCTION, a benchmark for finite structure concept synthesis in first order logic. Given small finite relational worlds with extensionally labeled target predicates,…
ReplaySCM: A Benchmark for Executable Causal Mechanism Induction from Interventions
Serafim Batzoglou
Most causal benchmarks for language models score local answers or graph structure. We introduce ReplaySCM, a 1,300 item benchmark for executable causal mechanism induction from fin…
ABD: Default Exception Abduction in Finite First Order Worlds
Serafim Batzoglou
We introduce ABD, a benchmark for default-exception abduction over finite first-order worlds. Given a background theory with an abnormality predicate and a set of relational struct…
Stress-Testing the Reasoning Competence of LLMs With Proofs Under Minimal Formalism
Konstantine Arkoudas, Serafim Batzoglou
We introduce ProofGrid, a benchmark suite for evaluating LLM reasoning through machine-checkable proofs rather than final answers alone. ProofGrid contains 15 tasks spanning proof…