3 papers
cs.CL2026
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
Jiadong Zhang, Xiaosong Ma
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the com…
cs.LG2025
LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long Reasoning
Haoyue Zhang, Hualei Zhang, Xiaosong Ma +2
Large Language Models (LLMs) exhibit enhanced capabilities by Chain-of-Thought reasoning. However, the extended reasoning sequences introduce significant GPU memory overhead due to…
cs.AI2025
Think How to Think: Mitigating Overthinking with Autonomous Difficulty Cognition in Large Reasoning Models
Yongjiang Liu, Haoxi Li, Xiaosong Ma +2
Recent Large Reasoning Models (LRMs) excel at complex reasoning tasks but often suffer from overthinking, generating overly long and redundant reasoning trajectories. To explore it…