3 papers
cs.LG2026
Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving
Hyungmin Kim, Minsoo Kim, Hongseok Kim +1
Multi-turn LLM serving accumulates dialogue history whose Key-Value (KV) cache grows with every turn and every user, quickly exceeding the model weights themselves and making memor…
cs.RO2025
Saliency-Aware Quantized Imitation Learning for Efficient Robotic Control
Seongmin Park, Hyungmin Kim, Sangwoo Kim +5
Deep neural network (DNN)-based policy models, such as vision-language-action (VLA) models, excel at automating complex decision-making from multi-modal inputs. However, scaling th…
cs.RO2024
Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control
Seongmin Park, Hyungmin Kim, Wonseok Jeon +4
Deep neural network (DNN)-based policy models like vision-language-action (VLA) models are transformative in automating complex decision-making across applications by interpreting…