activity
20242026
collaborators

9 papers

cs.AI2026

SETA: Scaling Environments for Terminal Agents

Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru +19

Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the te…

cs.AI2026

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

Dmitrii Khizbullin, Zaid Alyafeai, Abdelrahman Eldesokey +4

We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human…

cs.CV2026

MedCTA: A Benchmark for Clinical Tool Agents

Tajamul Ashraf, Hyewon Jeong, Fida Mohammad Thoker +1

To make clinically grounded decisions, medical AI agents are expected to go beyond simple recognition and be capable of tool retrieval, evidence acquisition, and integration. Exist…

cs.AI2026

AVA: Attentive VLM Agent for Mastering StarCraft II

Weiyu Ma, Yuqian Fu, Zecheng Zhang +2

We introduce AVACraft, a multimodal StarCraft II benchmark supporting both Multi-Agent Reinforcement Learning (MARL) and Vision-Language Model (VLM) paradigms. Unlike SMAC-family e…

cs.LG2025

Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers

Xingyue Huang, Rishabh, Gregor Franke +43

Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RL…

cs.AI2025

CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents

Tianqi Xu, Linyao Chen, Dai-Jie Wu +13

The development of autonomous agents increasingly relies on Multimodal Language Models (MLMs) to perform tasks described in natural language with GUI environments, such as websites…