2 papers
cs.AR2025
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
Dowon Kim, MinJae Lee, Janghyeon Kim +8
The expansion of context windows in large language models (LLMs) to multi-million tokens introduces severe memory and compute bottlenecks, particularly in managing the growing Key-…
q-bio.NC2025
Swin fMRI Transformer Predicts Early Neurodevelopmental Outcomes from Neonatal fMRI
Patrick Styll, Dowon Kim, Jiook Cha
Brain development in the first few months of human life is a critical phase characterized by rapid structural growth and functional organization. Accurately predicting developmenta…