3 papers
cs.CL2026
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
Zexue He, Yu Wang, Churan Zhi +11
Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversat…
cs.LG2025
Data Shifts Hurt CoT: A Theoretical Study
Lang Yin, Debangshu Banerjee, Gagandeep Singh
Chain of Thought (CoT) has been applied to various large language models (LLMs) and proven to be effective in improving the quality of outputs. In recent studies, transformers are…
cs.AI2024
On the Expressive Power of Tree-Structured Probabilistic Circuits
Lang Yin, Han Zhao
Probabilistic circuits (PCs) have emerged as a powerful framework to compactly represent probability distributions for efficient and exact probabilistic inference. It has been show…