8 citations · 11 across the 5 of their papers we have counts for
4 papers · 1 filter
Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures
Liam Cooper, Shinnung Jeong, Hyeran Jeon +2
Large language model (LLM) outputs are expected to be reproducible under greedy decoding, yet in practice the same model, prompt, and software stack produce different outputs on di…
Online Model Swapping in Architectural Simulation
Patrick Lavin, Jeffrey Young, Rich Vuduc +1
As systems and applications grow more complex, detailed simulation takes an ever increasing amount of time. The prospect of increased simulation time resulting in slower design ite…
Wrangling Rogues: Managing Experimental Post-Moore Architectures
Will Powell, Jason Riedy, Jeffrey S. Young +1
The Rogues Gallery is a new experimental testbed that is focused on tackling "rogue" architectures for the Post-Moore era of computing. While some of these devices have roots in th…
Performance Implications of NoCs on 3D-Stacked Memories: Insights from the Hybrid Memory Cube
Ramyad Hadidi, Bahar Asgari, Jeffrey Young +4
Memories that exploit three-dimensional (3D)-stacking technology, which integrate memory and logic dies in a single stack, are becoming popular. These memories, such as Hybrid Memo…