6 papers
Defeat the Heap: Zero-Copy Data Movement in AXI4MLIR
Elam Cohavi, Nicolas Bohm Agostini, Jude Haris +3
As custom hardware accelerators become increasingly central to machine learning workloads, efficient data transfer is critical for maximizing accelerator performance on linear alge…
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs
Rappy Saha, Jude Haris, Nicolas Bohm Agostini +2
Power-of-two (PoT) quantization significantly reduces the size of deep neural networks (DNNs) and replaces multiplications with bit-shift operations for inference. Prior work has s…
GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs
Lara D'Agata, Carlos Agulló-Domingo, Ãscar Vera-López +7
Fully homomorphic encryption (FHE) has recently attracted significant attention as both a cryptographic primitive and a systems challenge. Given the latest advances in accelerated…
FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption
Lohit Daksha, Seyda Guzelhan, Kaustubh Shivdikar +10
Fully Homomorphic Encryption (FHE) enables computation directly on encrypted data but incurs massive computational and memory overheads, often exceeding plaintext execution by seve…
Characterizing the Behavior of Training Mamba-based State Space Models on GPUs
Trinayan Baruah, Kaustubh Shivdikar, Sara Prescott +1
Mamba-based State Space Models (SSM) have emerged as a promising alternative to the ubiquitous transformers. Despite the expressive power of transformers, the quadratic complexity…
FIDESlib: A Fully-Fledged Open-Source FHE Library for Efficient CKKS on GPUs
Carlos Agulló-Domingo, Ãscar Vera-López, Seyda Guzelhan +7
Word-wise Fully Homomorphic Encryption (FHE) schemes, such as CKKS, are gaining significant traction due to their ability to provide post-quantum-resistant, privacy-preserving appr…