From the 1 of 4 linked papers with an AI index.
4 papers
A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM Inference
Jing Ding, Yash Nishant, Chandrish Ambati +2
The paper proposes a photonic‑CXL memory appliance that uses a passive fiber shuffle to provide a 32 TB shared memory pool for large‑language‑model KV cache management, reducing la…
Dissecting Embedding Bag Performance in DLRM Inference
Chandrish Ambati, Jing Ding, Trung Diep
As the size of DLRMs gets larger, the models must be partitioned across multiple GPUs or nodes of GPUs due to the size limitation of total HBM memory that can be packaged in a GPU.…
AMD MI300X GPU Performance Analysis
Chandrish Ambati, Trung Diep
The rapid growth of large language models (LLMs) has driven the need for high-performance, scalable GPU hardware capable of efficiently serving models with hundreds of billions of…
Photonic Fabric Platform for AI Accelerators
Jing Ding, Trung Diep
This paper presents the Photonic FabricTM and the Photonic Fabric ApplianceTM (PFA), a photonic-enabled switch and memory subsystem that delivers low latency, high bandwidth, and l…