◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Qichen Fu

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author2

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.CL2
  • cs.LG1
ORCID 0009-0008-5976-4095

identity via Semantic Scholar / OpenAlex

most citedLazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

1 citations · 1 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2024★ 1 cited

LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference

Qichen Fu, Minsik Cho, Thomas Merth +3

The inference of transformer-based large language models consists of two sequential stages: 1) a prefilling stage to compute the KV cache of prompts and generate the first token, a…

cs.CL2024

Speculative Streaming: Fast LLM Inference without Auxiliary Models

Nikhil Bhendawade, Irina Belousova, Qichen Fu +3

Speculative decoding is a prominent technique to speed up the inference of a large target language model based on predictions of an auxiliary draft model. While effective, in appli…

cs.LG2023

eDKM: An Efficient and Accurate Train-time Weight Clustering for Large Language Models

Minsik Cho, Keivan A. Vahid, Qichen Fu +5

Since Large Language Models or LLMs have demonstrated high-quality performance on many complex language tasks, there is a great interest in bringing these LLMs to mobile devices fo…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.