4 papers
Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs
Mohanad Odema, Gabrielle De Micheli, Dayin Gou +3
Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-le…
Lightweight, Practical Encrypted Face Recognition with GPU Support
Gabrielle De Micheli, Syed Mahbub Hafiz, Geovandro Pereira +5
Face recognition models operate in a client-server setting where a client extracts a compact face embedding and a server performs similarity search over a template database. This r…
HE-LRM: Encrypted Deep Learning Recommendation Models using Fully Homomorphic Encryption
Karthik Garimella, Austin Ebel, Gabrielle De Micheli +1
Fully Homomorphic Encryption (FHE) enables computation directly on encrypted data and privacy-preserving neural inference in the cloud. Existing solutions focus on models with dens…
CARVQ: Corrective Adaptor with Group Residual Vector Quantization for LLM Embedding Compression
Dayin Gou, Sanghyun Byun, Nilesh Malpeddi +4
Large Language Models (LLMs) typically rely on a large number of parameters for token embedding, leading to substantial storage requirements and memory footprints. In particular, L…