3 papers
cs.AI2026
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specif…
cs.CL2026
AdaMem: Learning What to Remember for Personalized Long-Horizon LLM Agents
Xingyu Chen, Rui Wang, Zhaopeng Tu +1
Long-term memory systems for Large Language Model (LLM) agents typically try to \emph{remember everything}, extracting memories uniformly to retain as many facts as possible. In pr…
cs.CL2025
Draft Model Knows When to Stop: Self-Verification Speculative Decoding for Long-Form Generation
Ziyin Zhang, Jiahao Xu, Tian Liang +4
Conventional speculative decoding (SD) methods utilize a predefined length policy for proposing drafts, which implies the premise that the target model smoothly accepts the propose…