most citedEvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

1 citations · 1 across the 6 of their papers we have counts for

collaborators

12 papers

cs.AI2026

DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models

ZhiYan Hou, Xinyu Tang, Hongyan An +9

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically verifiable outcome signals, but these signals…

cs.CV2026

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models

Yu Chen, Xiaohong Li, Xiaole Wang +3

In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply.…

cs.CL2026

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

Yanze Wang, Pengfei Yao, Tianyi Sun +7

Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Public repositories now host them i…

cs.CL2026

HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution

Xiaotian Luo, Fengxingyu Wang, Chuanrui Hu +2

Large Language Models (LLMs) have enabled capable agents across diverse applications. Beyond the foundation model, the performance of an agent is governed by the surrounding agent…

cs.AI20261 cited

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Xingze Gao, Chuanrui Hu, Hongda Chen +9

Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and ve…

cs.AI2026

LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis

Shihao Xu, Tiancheng Zhou, Jiatong Ma +8

Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely…