3 papers
cs.OS2026
mzCache: On-Device LLM Memory Management under Multitasking
Hongseung Yu, Minsung Kim, Jongseok Park +1
On-device mobile Large Language Model (LLM) inference is gaining significant attention. However, mobile devices operate in highly dynamic multitasking environments where users freq…
cs.LG2025
SBVR: Summation of BitVector Representation for Efficient LLM Quantization
Wonjun Bang, Jongseok Park, Hongseung Yu +2
With the advent of large language models (LLMs), numerous Post-Training Quantization (PTQ) strategies have been proposed to alleviate deployment barriers created by their enormous…
cs.NI2024
Towards a Dynamic Future with Adaptable Computing and Network Convergence (ACNC)
Masoud Shokrnezhad, Hao Yu, Tarik Taleb +4
In the context of advancing 6G, a substantial paradigm shift is anticipated, highlighting comprehensive everything-to-everything interactions characterized by numerous connections…