5 papers
Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey +4
We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human…
Beyond Outlining: Heterogeneous Recursive Planning for Adaptive Long-form Writing with Language Models
Ruibin Xiong, Yimeng Chen, Dmitrii Khizbullin +2
Long-form writing agents require flexible integration and interaction across information retrieval, reasoning, and composition. Current approaches rely on predefined workflows and…
How to Correctly do Semantic Backpropagation on Language-based Agentic Systems
Wenyi Wang, Hisham A. Alyahya, Dylan R. Ashley +4
Language-based agentic systems have shown great promise in recent years, transitioning from solving small-scale research problems to being deployed in challenging real-world tasks.…
Agent-as-a-Judge: Evaluate Agents with Agents
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley +10
Contemporary evaluation techniques are inadequate for agentic systems. These approaches either focus exclusively on final outcomes -- ignoring the step-by-step nature of agentic sy…
Language Agents as Optimizable Graphs
Mingchen Zhuge, Wenyi Wang, Louis Kirsch +3
Various human-designed prompt engineering techniques have been proposed to improve problem solvers based on Large Language Models (LLMs), yielding many disparate code bases. We uni…