Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SETA: Scaling Environments for Terminal Agents
Qijia Shen, Zhiqi Huang, Vamsidhar Kamanuru +19
Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the te…
cs.AI2026
X-RAY: Mapping LLM Reasoning Capability via Formalized and Calibrated Probes
Tianxi Gao, Yufan Cai, Yusi Yuan +1
Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood. Existing evaluations largely emphasize task-level accuracy, often…