◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Kuntai Du

15 papers hereh-index 151.1k citations36 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author12

Across the 14 of 15 papers where every author was matched, so the position is known.

fields
  • cs.LG5
  • cs.AI2
  • cs.DC2
  • cs.OS2
  • cs.AR1
  • cs.CL1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.OSShow all

2 papers · 1 filter

cs.OS2026

AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving

Shaoting Feng, Hanchen Li, Kuntai Du +8

Large language model (LLM) applications often reuse previously processed context, such as chat history and documents, which introduces significant redundant computation. Existing L…

cs.OS2025

EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving

Shaoting Feng, Yuhan Liu, Hanchen Li +11

Reusing KV cache is essential for high efficiency of Large Language Model (LLM) inference systems. With more LLM users, the KV cache footprint can easily exceed GPU memory capacity…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.