2 papers
cs.CL2026
Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
Fali Wang, Ali Al-Lawati, Iliyas Bektas +5
Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurall…
cs.AI2025
NS-Gym: Open-Source Simulation Environments and Benchmarks for Non-Stationary Markov Decision Processes
Nathaniel S. Keplinger, Baiting Luo, Iliyas Bektas +5
In many real-world applications, agents must make sequential decisions in environments where conditions are subject to change due to various exogenous factors. These non-stationary…