3 papers
cs.CL2026
MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation
Kuan Wang
Agent memory systems are increasingly evaluated against RAG and full-context baselines, but reported gains often mix changes in the memory method with changes in the language model…
cs.LG2026
MemLeak: Diagnosing Information Leaks in Multimodal Agent Memory
Kuan Wang, Chao Zhang
When a multimodal AI agent is asked to forget a fact, current memory systems usually delete the text entry and report success. We find that the fact can remain recoverable from ret…
cs.RO2025
What Matters in Learning from Large-Scale Datasets for Robot Manipulation
Vaibhav Saxena, Matthew Bronars, Nadun Ranawaka Arachchige +5
Imitation learning from large multi-task demonstration datasets has emerged as a promising path for building generally-capable robots. As a result, 1000s of hours have been spent o…