2 citations · 3 across the 6 of their papers we have counts for
7 papers
Measuring AI Reasoning: A Guide for Researchers
Munachiso Samuel Nwadike, Zangir Iklassov, Kareem Ali +2
In this paper, we offer a guide for researchers on evaluating reasoning in language models, building the case that reasoning should be assessed through evidence of adaptive, multi-…
The AI Data Scientist
Farkhad Akimov, Munachiso Samuel Nwadike, Zangir Iklassov +1
Imagine decision-makers uploading data and, within minutes, receiving clear, actionable insights delivered straight to their fingertips. That is the promise of the AI Data Scientis…
SVRPBench: A Realistic Benchmark for Stochastic Vehicle Routing Problem
Ahmed Heakl, Yahia Salaheldin Shaaban, Martin Takac +2
Robust routing under uncertainty is central to real-world logistics, yet most benchmarks assume static, idealized settings. We present SVRPBench, the first open benchmark to captur…
LLM-BABYBENCH: Understanding and Evaluating Grounded Planning and Reasoning in LLMs
Omar Choukrani, Idriss Malek, Daniil Orel +4
Assessing the capacity of Large Language Models (LLMs) to plan and reason within the constraints of interactive environments is crucial for developing capable AI agents. We introdu…
RECALL: Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal Cycles
Munachiso Nwadike, Zangir Iklassov, Toluwani Aremu +6
We introduce the concept of the self-referencing causal cycle (abbreviated RECALL) - a mechanism that enables large language models (LLMs) to bypass the limitations of unidirection…
A Decade of Deep Learning: A Survey on The Magnificent Seven
Dilshod Azizov, Muhammad Arslan Manzoor, Velibor Bojkovic +9
Deep learning has fundamentally reshaped the landscape of artificial intelligence over the past decade, enabling remarkable achievements across diverse domains. At the heart of the…