3 papers
cs.LG2026
VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG
Junkyum Kim, Divya Mahajan
Retrieval-Augmented Generation (RAG) systems combine vector similarity search with large language models (LLMs) to deliver accurate, context-aware responses. However, co-locating t…
cs.AR2025
Flexible In-NAND Cryptographic Processing for Secure Flash Storage
Seock-Hwan Noh, Hoyeon Lee, Junkyum Kim +6
We present FlashVault, an in-NAND self-encryption architecture that embeds a reconfigurable cryptographic engine into the unused silicon area of a state-of-the-art 4D V-NAND struct…
cs.DC2025
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
Seonho Lee, Jihwan Oh, Junkyum Kim +3
This paper provides an in-depth characterization of GPU-accelerated systems, to understand the interplay between overlapping computation and communication which is commonly employe…