2 citations · 2 across the 1 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
Yeonjoon Jung, Daehyun Ahn, Hyungjun Kim +2
Low-Rank Adaptation (LoRA) is a popular method for parameter-efficient fine-tuning (PEFT) of generative models, valued for its simplicity and effectiveness. Despite recent enhancem…
cs.LG2024★ 2 cited
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
Taesu Kim, Jongho Lee, Daehyun Ahn +4
We introduce QUICK, a group of novel optimized CUDA kernels for the efficient inference of quantized Large Language Models (LLMs). QUICK addresses the shared memory bank-conflict p…
cs.LG2023
Squeezing Large-Scale Diffusion Models for Mobile
Jiwoong Choi, Minkyu Kim, Daehyun Ahn +6
The emergence of diffusion models has greatly broadened the scope of high-fidelity image synthesis, resulting in notable advancements in both practical implementation and academic…