Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models
Tianlong Wang, Pinqiao Wang, Weili Shi +1
Large language models (LLMs) with advanced cognitive capabilities are emerging as agents for various reasoning and planning tasks. Traditional evaluations often focus on specific r…
cs.AI2025
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
Guangya Wan, Mingyang Ling, Xiaoqi Ren +3
Long-horizon tasks that require sustained reasoning and multiple tool interactions remain challenging for LLM agents: small errors compound across steps, and even state-of-the-art…
cs.AI2025
Memory in Large Language Models: Mechanisms, Evaluation and Evolution
Dianxing Zhang, Wendong Li, Kani Song +4
Under a unified operational definition, we define LLM memory as a persistent state written during pretraining, finetuning, or inference that can later be addressed and that stably…