collaborators

6 papers

cs.MA2026

Muscle Memory for Agents: Compile not Merely Retrieve

Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang +1

Memory for LLM agents has converged on a single architectural pattern: store experience as text, embeddings, reflections, or rules; retrieve at inference time; let a general-purpos…

cs.CV2026

Thinking with Anchors: Grounded and Efficient Document Reasoning

Sichen Zhu, Yuchen Zhu, Wenzhuo Xu +13

Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reason jointly about region seman…

cs.CV2026

STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models

Yiming Liang, Yixiao Chen, Yiyang Zhou +8

Many video reasoning tasks require tracking motion, temporal order, and evolving visual states across frames. Existing methods built on large vision-language models (LVLMs) often a…

cs.MM2026

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval

Yawen Qin, Ke Qiu, Qin Zhang

Dance serves as both a cultural cornerstone and a medium for personal expression, yet the rapid growth of online dance content has made personalized discovery increasingly difficul…

cs.DC2025

DeepServe: Serverless Large Language Model Serving at Scale

Junhao Hu, Jiang Xu, Zhixia Liu +18

In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addr…

cs.LG2025

EPIC: Efficient Position-Independent Caching for Serving Large Language Models

Junhao Hu, Wenrui Huang, Weidong Wang +7

Large Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become mor…