17 papers
High-Performance NTT Accelerators for PQC leveraging Unified Redundant Arithmetic and Fine-Tuned Microarchitecture
George Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos
Post-quantum cryptography and privacy-preserving technologies are expected to play a central role in future secure communication systems. Lattice-based PQC schemes such as ML-KEM (…
Low-Cost Multi-Precision Systolic Arrays for Accelerating FHE NTTs on AI ASICs
George Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos
Fully Homomorphic Encryption (FHE) ensures robust data privacy but suffers from prohibitive computational overhead. Accelerating FHE on AI hardware like Tensor Processing Units (TP…
MIVE: A Minimalist Integer Vector Engine for Softmax LayerNorm and RMSNorm Acceleration
Kosmas Alexandridis, Giorgos Dimitrakopoulos
The rapid growth of Large Language Models (LLMs) has intensified the need for specialized hardware accelerators that can satisfy stringent inference latency and power constraints.…
MPX: A Unified Systolic Array for Matrix and Polynomial Multiplication
George Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos
Polynomial multiplication is a fundamental kernel in Fully Homomorphic Encryption (FHE) and post-quantum cryptography (PQC) and is commonly accelerated through Number Theoretic Tra…
VQ4SNN: Vector Quantization for Memory-Efficient FPGA Spiking Neural Networks
Dimitrios Sekertzis, Giorgos Dimitrakopoulos
Spiking Neural Networks (SNNs) offer an energy-efficient paradigm for edge AI, making them attractive for hardware acceleration. However, deploying dense SNNs on FPGAs is constrain…
H-FA: A Hybrid Floating-Point and Logarithmic Approach to Hardware Accelerated FlashAttention
Kosmas Alexandridis, Giorgos Dimitrakopoulos
Transformers have significantly advanced AI and machine learning through their powerful attention mechanism. However, computing attention on long sequences can become a computation…