2 papers
cs.AI2025
Shared Disk KV Cache Management for Efficient Multi-Instance Inference in RAG-Powered LLMs
Hyungwoo Lee, Kihyun Kim, Jinwoo Kim +5
Recent large language models (LLMs) face increasing inference latency as input context length and model size continue to grow. In particular, the retrieval-augmented generation (RA…
physics.flu-dyn2024
Influence of three-dimensionality on wake synchronization of oscillatory cylinder
Youngjae Kim, Vedasri Godavarthi, Laura Victoria Rolandi +2
We investigate the effect of three-dimensionality on the synchronization characteristics of the wake behind an oscillating circular cylinder at Re = 300. Cylinder oscillations in r…