collaborators

5 papers

cs.LG2026

SplitLite: Low-Rank Residual Compression for Split Learning

Tao Li, Yulin Tang, Qi Guo +1

Federated fine-tuning of on-device large language models (LLMs) faces a significant computing burden. To overcome this limitation, split learning (SL) has emerged as a promising so…

cs.NI2026

SplitCom: Communication-efficient Split Federated Fine-tuning of LLMs via Temporal Compression

Tao Li, Yulin Tang, Yiyang Song +4

Federated fine-tuning of on-device large language models (LLMs) mitigates privacy concerns by preventing raw data sharing. However, the intensive computational and memory demands p…

cs.NI2025

SLIDE: Simultaneous Model Downloading and Inference at the Wireless Network Edge

Guanqiao Qu, Tao Li, Qian Chen +2

To support on-device inference, the next-generation mobile networks are expected to support real-time model downloading services to mobile users. However, powerful AI models typica…

cs.CR2025

NWaaS: A Non-Intrusive and Privacy-Preserving Watermarking-as-a-Service System with Adaptive Resource Scheduling

Haonan An, Qianyao Ren, Guang Hua +5

Securing intellectual property (IP) in Machine Learning as a Service is critical yet challenging. While deep neural network watermarking serves as a standard defense against model…

cs.DC2025

Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

Zonghang Li, Tao Li, Wenjiao Feng +8

On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability. To overcome th…