9 papers
SymStep: Symbolic Step Verification for Logical Reasoning
Aida Usmanova, Rui Gao, Dilshod Azizov +2
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an L…
Einstein World Models
Munachiso Samuel Nwadike, Zangir Iklassov, Ali Mekky +2
Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex thought cannot be captured through language alo…
Measuring AI Reasoning: A Guide for Researchers
Munachiso Samuel Nwadike, Zangir Iklassov, Kareem Ali +2
In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed through evidence of adaptive, multi-…
The AI Data Scientist
Farkhad Akimov, Munachiso Samuel Nwadike, Zangir Iklassov +1
Imagine decision-makers uploading data and, within minutes, receiving clear, actionable insights delivered straight to their fingertips. That is the promise of the AI Data Scientis…
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
Ahmed Heakl, Yahia Salaheldin Shaaban, Martin Takac +2
Robust routing under uncertainty is central to real-world logistics, yet most benchmarks assume static, idealized settings. We present SVRPBench, the first open benchmark to captur…
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Omar Choukrani, Idriss Malek, Daniil Orel +4
Assessing the capacity of Large Language Models (LLMs) to plan and reason within the constraints of interactive environments is crucial for developing capable AI agents. We introdu…