3 papers
cs.CL2025
A Survey on Inference Engines for Large Language Models: Perspectives on Optimization and Efficiency
Sihyeong Park, Sungryeol Jeon, Chaelyn Lee +3
Large language models (LLMs) are widely applied in chatbots, code generators, and search engines. Workload such as chain-of-throught, complex reasoning, agent services significantl…
cs.CV2025
IPTQ-ViT: Post-Training Quantization of Non-linear Functions for Integer-only Vision Transformers
Gihwan Kim, Jemin Lee, Hyungshin Kim
Previous Quantization-Aware Training (QAT) methods for vision transformers rely on expensive retraining to recover accuracy loss in non-linear layer quantization, limiting their us…
cs.PF2025
A Review on Proprietary Accelerators for Large Language Models
Sihyeong Park, Jemin Lee, Byung-Soo Kim +1
With the advancement of Large Language Models (LLMs), the importance of accelerators that efficiently process LLM computations has been increasing. This paper discusses the necessi…