3 papers
cs.SE2026
The Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent Behavior
Xiangzhe Xu, Hamidreza Saghir, Qianhui Wu +5
As large language models continue to improve, agentic systems are becoming increasingly important, and tools are a key design dimension because they determine how agents access inf…
cs.AI2026
MemWM: Memory-Augmented Text-Based World Model
Yujun Wang, Tao Zhang, Jinhe Bi +9
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can sti…
cs.CL2025
DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
Chiyu Zhang, Marc-Alexandre Cote, Michael Albada +6
Large language model (LLM) agents have shown impressive capabilities in human language comprehension and reasoning, yet their potential in cybersecurity remains underexplored. We i…