8 papers
SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents
Yanze Wang, Pengfei Yao, Tianyi Sun +7
Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Public repositories now host them i…
HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution
Xiaotian Luo, Fengxingyu Wang, Chuanrui Hu +2
Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is governed by the surrounding agent…
EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
Xingze Gao, Chuanrui Hu, Hongda Chen +9
Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and ve…
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Yu Chen, Runkai Chen, Sheng Yi +9
Long-term memory is a cornerstone of human intelligence. Enabling AI to process lifetime-scale information remains a long-standing pursuit in the field. Due to the constraints of f…
HyperMem: Hypergraph Memory for Long-Term Conversations
Juwei Yue, Chuanrui Hu, Jiawei Sheng +5
Long-term memory is essential for conversational agents to maintain coherence, track persistent tasks, and provide personalized interactions across extended dialogues. However, exi…
Evaluating Long-Horizon Memory for Multi-Party Collaborative Dialogues
Chuanrui Hu, Tong Li, Xingze Gao +8
Long-term conversational memory in practical LLM applications is inherently collaborative: information is produced by multiple participants, scattered across groups and channels, r…