◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Han Guo

8 papers hereh-index 8387 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author3
  • middle author5

Across the 8 of 8 papers where every author was matched, so the position is known.

fields
  • cs.LG5
  • cs.CL2
  • cs.AI1
same name
  • Han Guo — 13 papers, h 17
  • Han Guo — 6 papers, h 3
  • Han Guo — 4 papers
  • Han Guo — 4 papers, h 3
  • Han Guo — 3 papers, h 2
  • Han Guo — 3 papers, h 5

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Fast KV Compaction via Attention Matching

Adam Zweiger, Xinghong Fu, Han Guo +1

Scaling language models to long contexts is often bottlenecked by the size of the key-value (KV) cache. In deployed settings, long contexts are typically managed through compaction…

cs.LG2025

Self-Adapting Language Models

Adam Zweiger, Jyothish Pari, Han Guo +3

Large language models (LLMs) are powerful but static; they lack mechanisms to adapt their weights in response to new tasks, knowledge, or examples. We introduce Self-Adapting LLMs…

cs.LG2025

Log-Linear Attention

Han Guo, Songlin Yang, Tarushii Goel +3

The attention mechanism in Transformers is an important primitive for accurate and scalable sequence modeling. Its quadratic-compute and linear-memory complexity however remain sig…

cs.LG2025

On the Duality between Gradient Transformations and Adapters

Lucas Torroba-Hennigen, Hunter Lang, Han Guo +1

We study memory-efficient optimization of neural networks (in particular language models) with linear gradient transformations, where the gradients are linearly mapped to a lower d…

cs.LG2024

Fast Matrix Multiplications for Lookup Table-Quantized LLMs

Han Guo, William Brandon, Radostin Cholakov +3

The deployment of large language models (LLMs) is often constrained by memory bandwidth, where the primary bottleneck is the cost of transferring model parameters from the GPU's gl…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.