◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Minkyu Kim

4 papers hereh-index 235 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.CV1
same name
  • Minkyu Kim — 12 papers, h 13
  • Minkyu Kim — 7 papers, h 8
  • Minkyu Kim — 5 papers, h 5
  • Minkyu Kim — 3 papers, h 3
  • Minkyu Kim — 3 papers, h 1
  • Minkyu Kim — 3 papers, h 2

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedQUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference

2 citations · 2 across the 2 of their papers we have counts for

collaborators

4 papers

cs.LG2026

Locality-Aware Redundancy Pruning for LLM Depth Compression

Vincent-Daniel Yun, Youngrae Kim, Woosang Lim +3

Large language models are known to contain representational redundancy across network depth, making depth pruning an effective approach for improving inference efficiency. Existing…

cs.LG2025

Neural Weight Compression for Language Models

Jegwang Ryu, Minkyu Kim, Seungjun Shin +3

Efficient compression of language model weights is increasingly critical as model scale and deployment grow. Yet, most existing methods rely on handcrafted transforms and heuristic…

cs.CV2024

Neural Image Compression with Text-guided Encoding for both Pixel-level and Perceptual Fidelity

Hagyeong Lee, Minkyu Kim, Jun-Hyuk Kim +3

Recent advances in text-guided image compression have shown great potential to enhance the perceptual quality of reconstructed images. These methods, however, tend to have signific…

cs.LG2024★ 2 cited

QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference

Taesu Kim, Jongho Lee, Daehyun Ahn +4

We introduce QUICK, a group of novel optimized CUDA kernels for the efficient inference of quantized Large Language Models (LLMs). QUICK addresses the shared memory bank-conflict p…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.