Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko +1
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevant context across extended re…
cs.AI2025
What Limits Agentic Systems Efficiency?
Song Bian, Minghao Yan, Anand Jayarajan +2
Large Language Models (LLMs), such as OpenAI-o1 and DeepSeek-R1, have demonstrated strong reasoning capabilities. To further enhance LLM capabilities, recent agentic systems, such…
cs.AI2024
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
Honghua Dong, Qidong Su, Yubo Gao +5
Large Language Models (LLMs) have become increasingly capable of handling diverse tasks with the aid of well-crafted prompts and integration of external tools, but as task complexi…