◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Harry Dong

14 papers hereh-index 7283 citations14 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author9
  • middle author5

Across the 14 of 14 papers where every author was matched, so the position is known.

fields
  • cs.LG5
  • cs.AI3
  • cs.CV2
  • stat.ML2
  • cs.CL1
  • eess.IV1

identity via Semantic Scholar / OpenAlex

activity
20222026
most citedShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

1 citations · 2 across the 14 of their papers we have counts for

collaborators
Showing 2024 · cs.LGShow all

3 papers · 2 filters

cs.LG2024★ 1 cited

ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inference

Hanshi Sun, Li-Wen Chang, Wenlei Bao +6

With the widespread deployment of long-context large language models (LLMs), there has been a growing demand for efficient support of high-throughput inference. However, as the key…

cs.LG2024

Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation

Harry Dong, Beidi Chen, Yuejie Chi

With the development of transformer-based large language models (LLMs), they have been applied to many fields due to their remarkable utility, but this comes at a considerable comp…

cs.LG2024★ 1 cited

Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM Inference

Harry Dong, Xinyu Yang, Zhenyu Zhang +3

Many computational factors limit broader deployment of large language models. In this paper, we focus on a memory bottleneck imposed by the key-value (KV) cache, a computational sh…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.