3 papers
cs.CL2026
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Tianxin Wei, Noveen Sachdeva, Benjamin Coleman +12
Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critical component, yet its management and ev…
cs.AI2026
Agentic Reasoning for Large Language Models
Tianxin Wei, Ting-Wei Li, Zhining Liu +26
Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) demonstrate strong reasoning capabilitie…
cs.CL2026
CoReflect: Conversational Evaluation via Co-Evolutionary Simulation and Reflective Rubric Refinement
Yunzhe Li, Richie Yueqi Feng, Tianxin Wei +1
Evaluating conversational systems in multi-turn settings remains a fundamental challenge. Conventional pipelines typically rely on manually defined rubrics and fixed conversational…