2 papers
cs.AI2026
ContextWeave: A Real-World Workflow Benchmark
Bo Wang, Yuqian Yao, Enxi Wang +25
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We…
cs.CL2025
SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning
Renxi Wang, Honglin Mu, Liqun Ma +5
Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchma…