activity
20242026
collaborators

6 papers

cs.AI2026

MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory

Beidi Zhao, Yaoqi Chen, Yuru Feng +10

Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a query may depend on evidence buried far bac…

cs.AI2026

SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation

Yuru Feng, Yaoqi Chen, Beidi Zhao +7

Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalignment. Real-world deployments thus requir…

cs.AI2026

Beyond Semantic Organization: Memory as Execution State Management for Long-Horizon Agents

Yaoqi Chen, Haibin Lai, Yuru Feng +12

LLM-based agents increasingly tackle long-horizon tasks with interdependent decisions, where each action reshapes future constraints and intermediate errors can cascade. Existing R…

cs.DC2025

Scalable Distributed Vector Search via Accuracy Preserving Index Construction

Yuming Xu, Qianxi Zhang, Qi Chen +8

Scaling Approximate Nearest Neighbor Search (ANNS) to billions of vectors requires distributed indexes that balance accuracy, latency, and throughput. Yet existing index designs st…

cs.DC2025

DISTRIBUTEDANN: Efficient Scaling of a Single DISKANN Graph Across Thousands of Computers

Philip Adams, Menghao Li, Shi Zhang +6

We present DISTRIBUTEDANN, a distributed vector search service that makes it possible to search over a single 50 billion vector graph index spread across over a thousand machines t…

cs.IR2024

MS MARCO Web Search: a Large-scale Information-rich Web Dataset with Millions of Real Click Labels

Qi Chen, Xiubo Geng, Corby Rosset +28

Recent breakthroughs in large models have highlighted the critical significance of data scale, labels and modals. In this paper, we introduce MS MARCO Web Search, the first large-s…