activity
20172022
most citedSCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks

125 citations · 289 across the 11 of their papers we have counts for

collaborators

12 papers

cs.LG20221 cited

HEAT: Hardware-Efficient Automatic Tensor Decomposition for Transformer Compression

Jiaqi Gu, Ben Keller, Jean Kossaifi +3

Transformers have attained superior performance in natural language processing and computer vision. Their self-attention and feedforward layers are overparameterized, limiting infe…

cs.LG20222 cited

An Adversarial Active Sampling-based Data Augmentation Framework for Manufacturable Chip Design

Mingjie Liu, Haoyu Yang, Zongyi Li +7

Lithography modeling is a crucial problem in chip design to ensure a chip design mask is manufacturable. It requires rigorous simulations of optical and chemical models that are co…

cs.OH2022

Generic Lithography Modeling with Dual-band Optics-Inspired Neural Networks

Haoyu Yang, Zongyi Li, Kumara Sastry +6

Lithography simulation is a critical step in VLSI design and optimization for manufacturability. Existing solutions for highly accurate lithography simulation with rigorous models…

cs.LG20221 cited

GATSPI: GPU Accelerated Gate-Level Simulation for Power Improvement

Yanqing Zhang, Haoxing Ren, Akshay Sridharan +1

In this paper, we present GATSPI, a novel GPU accelerated logic gate simulator that enables ultra-fast power estimation for industry sized ASIC designs with millions of gates. GATS…

cs.LG20211 cited

NVCell: Standard Cell Layout in Advanced Technology Nodes with Reinforcement Learning

Haoxing Ren, Matthew Fojtik, Brucek Khailany

High quality standard cell layout automation in advanced technology nodes is still challenging in the industry today because of complex design rules. In this paper we introduce an…

cs.AR2021

Softermax: Hardware/Software Co-Design of an Efficient Softmax for Transformers

Jacob R. Stevens, Rangharajan Venkatesan, Steve Dai +2

Transformers have transformed the field of natural language processing. This performance is largely attributed to the use of stacked self-attention layers, each of which consists o…