3 papers
cs.CR2026
ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
Jianhang Xie, Sicheng Tan, Vishnu Naresh Boddeti +1
Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under FHE remains prohibitively expensive. A key…
cs.LG2025
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
Jianhang Xie, Chuntao Ding, Xiaqing Li +3
Deploying quantized deep neural network (DNN) models with resource adaptation capabilities on ubiquitous Internet of Things (IoT) devices to provide high-quality AI services can le…
cs.DC2024
LoRA-C: Parameter-Efficient Fine-Tuning of Robust CNN for IoT Devices
Chuntao Ding, Xu Cao, Jianhang Xie +3
Efficient fine-tuning of pre-trained convolutional neural network (CNN) models using local data is essential for providing high-quality services to users using ubiquitous and resou…