3 papers
cs.DC2026
Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence
Chanwoo Cho, Wooseok Kim, Yonglak Son +2
Large language models (LLMs) are widely used in intelligent services due to their remarkable capability in generative tasks. Typically, LLM-based services process the inference req…
cs.CV2025
CorGi: Contribution-Guided Block-Wise Interval Caching for Training-Free Acceleration of Diffusion Transformers
Yonglak Son, Suhyeok Kim, Seungryong Kim +1
Diffusion transformer (DiT) achieves remarkable performance in visual generation, but its iterative denoising process combined with larger capacity leads to a high inference cost.…
cs.CL2025
Not All Adapters Matter: Selective Adapter Freezing for Memory-Efficient Fine-Tuning of Language Models
Hyegang Son, Yonglak Son, Changhoon Kim +1
Transformer-based large-scale pre-trained models achieve great success. Fine-tuning is the standard practice for leveraging these models in downstream tasks. Among the fine-tuning…