5 papers
Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift
Weihang Su, Jiacheng Kang, Jingyan Xu +7
Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumulate weight drift, potentially causing cat…
Skill Retrieval Augmentation for Agentic AI
Weihang Su, Jianming Long, Qingyao Ai +6
As large language models (LLMs) evolve into agentic problem solvers, they increasingly rely on external, reusable skills to handle tasks beyond their native parametric capabilities…
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
Qingyao Ai, Yichen Tang, Changyue Wang +3
Scaling up data, parameters, and test-time computation has been the mainstream methods to improve LLM systems (LLMsys), but their upper bounds are almost reached due to the gradual…
SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
Weihang Su, Anzhe Xie, Qingyao Ai +5
The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this proces…
Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
Weihang Su, Jianming Long, Changyue Wang +5
Large Language Models (LLMs) frequently exhibit hallucinations, generating content that appears fluent and coherent but is factually incorrect. Such errors undermine trust and hind…