4 papers · 1 filter
From Rollouts to Recipes: Self-Contained Post-Training for LLMs
Yifei Li, Lingling Zhang, Muye Huang +3
Post-training large language models usually applies a single training recipe to all samples, even though the model's own rollouts reveal different sample-level learning states. We…
QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents
Heng Wang, Yifei Li, Lingling Zhang +4
Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user preferenc…
GeoChallenge: A Multi-Answer Multiple-Choice Benchmark for Geometric Reasoning with Diagrams
Yushun Zhang, Weiping Fu, Zesheng Yang +6
Evaluating the symbolic reasoning of large language models (LLMs) calls for geometry benchmarks that require multi-step proofs grounded in both text and diagrams. However, existing…
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
Yifei Li, Weidong Guo, Lingling Zhang +6
Long-term conversational memory is a core capability for LLM-based dialogue systems, yet existing benchmarks and evaluation protocols primarily focus on surface-level factual recal…