Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ContextGuard: Structured Self-Auditing for Context Learning in Language Models
Hongbo Jin, Chi Wang, Haoran Tang +5
Recent benchmarks reveal that despite strong reasoning capabilities, large language models (LLMs) still struggle to faithfully apply complex contextual knowledge. These failures ar…
cs.CL2026
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Tianxin Wei, Noveen Sachdeva, Benjamin Coleman +12
Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critical component, yet its management and ev…