1 citations · 1 across the 2 of their papers we have counts for
4 papers
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
Zonghang Li, Tao Li, Wenjiao Feng +8
On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability. To overcome th…
SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge
Guanqiao Qu, Tao Li, Qian Chen +2
To support on-device inference, the next-generation mobile networks are expected to support real-time model downloading services to mobile users. However, powerful AI models typica…
SplitCom: Communication-efficient Split Federated Fine-tuning of LLMs via Temporal Compression
Tao Li, Yulin Tang, Yiyang Song +4
Federated fine-tuning of on-device large language models (LLMs) mitigates privacy concerns by preventing raw data sharing. However, the intensive computational and memory demands p…
NWaaS: A Non-Intrusive and Privacy-Preserving Watermarking-as-a-Service System with Adaptive Resource Scheduling
Haonan An, Guang Hua, Qianyao Ren +6
Securing intellectual property (IP) in Machine Learning as a Service is critical yet challenging. While deep neural network watermarking serves as a standard defense against model…