3 citations · 3 across the 3 of their papers we have counts for
3 papers
GPU Domain Specialization via Composable On-Package Architecture
Yaosheng Fu, Evgeny Bolotin, Niladrish Chatterjee +2
As GPUs scale their low precision matrix math throughput to boost deep learning (DL) performance, they upset the balance between math throughput and memory system capabilities. We…
The Architectural Implications of Distributed Reinforcement Learning on CPU-GPU Systems
Ahmet Inci, Evgeny Bolotin, Yaosheng Fu +4
With deep reinforcement learning (RL) methods achieving results that exceed human capabilities in games, robotics, and simulated environments, continued scaling of RL training is c…
Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUs
Esha Choukse, Michael Sullivan, Mike O'Connor +4
GPUs offer orders-of-magnitude higher memory bandwidth than traditional CPU-only systems. However, GPU device memory tends to be relatively small and the memory capacity can not be…