most citedRetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.AI2026

MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory

Beidi Zhao, Yaoqi Chen, Yuru Feng +10

Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far bac…

cs.AI2026

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

Yuru Feng, Yaoqi Chen, Beidi Zhao +7

Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus requir…

cs.CL2026

TriggerBench: Investigating Prospective Memory for Large Language Models

Tianhua Zhang, Xinjiang Wang, Qianxi Zhang +6

While Large Language Models (LLMs) are increasingly deployed in long interactions, existing evaluations focus predominantly on retrospective memory (RM) via explicit queries. Prosp…

cs.AI2026

Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents

Yaoqi Chen, Haibin Lai, Yuru Feng +12

LLM-based agents increasingly tackle long-horizon tasks with interdependent decisions, where each action reshapes future constraints and intermediate errors can cascade. Existing R…

cs.LG20261 cited

RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference

Yaoqi Chen, Jinkai Zhang, Baotong Lu +16

Recent large language models (LLMs) are rapidly extending their context windows, yet inference throughput lags due to increasing GPU memory and bandwidth demands. This is because t…