4 papers
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges
Dowon Son, Yonggon Park, Hyunuk Cho +5
This work investigates the potential benefits and technical challenges of using high-bandwidth flash (HBF) for large language model (LLM) inference. HBF has gained increasing atten…
Flexible In-NAND Cryptographic Processing for Secure Flash Storage
Seock-Hwan Noh, Hoyeon Lee, Junkyum Kim +6
We present FlashVault, an in-NAND self-encryption architecture that embeds a reconfigurable cryptographic engine into the unused silicon area of a state-of-the-art 4D V-NAND struct…
A Diffusion-Based Framework for Configurable and Realistic Multi-Storage Trace Generation
Seohyun Kim, Junyoung Lee, Jongho Park +3
We propose DiTTO, a novel diffusion-based framework for generating realistic, precisely configurable, and diverse multi-device storage traces. Leveraging advanced diffusion techniq…
Rethinking Caching for LLM Serving Systems: Beyond Traditional Heuristics
Jungwoo Kim, Minsang Kim, Jaeheon Lee +6
Serving Large Language Models (LLMs) at scale requires meeting strict Service Level Objectives (SLOs) under severe computational and memory constraints. Nevertheless, traditional c…