3 papers
cs.AI2026
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We ask whether LLMs can instead discover one another's weaknesses and turn those o…
cs.CL2026
AdaMem: Learning What to Remember with Adaptive Memory Policies for Personalized Agents
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Long-term memory systems allow LLM agents to preserve information beyond a single context window, but most systems focus on storing and retrieving facts after extraction, leaving t…
cs.CL2024
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
Ziyin Zhang, Jiahao Xu, Tian Liang +4
Conventional speculative decoding (SD) methods utilize a predefined length policy for proposing drafts, which implies the premise that the target model smoothly accepts the propose…