From the 1 of 16 linked papers with an AI index.
3 citations · 3 across the 9 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey +4
We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human…
cs.AI2026
FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured Representations
Fedor Rodionov, Abdelrahman Eldesokey, Michael Birsak +3
We introduce FloorplanQA, a diagnostic benchmark for evaluating spatial reasoning in large language models (LLMs). FloorplanQA is grounded in structured representations of indoor s…