activity
20242026
most citedPrima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.DC20261 cited

Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters

Zonghang Li, Tao Li, Wenjiao Feng +8

On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability. To overcome th…

cs.NI2026

A Fragmentation-Aware Adaptive Bilevel Search Framework for Service Mapping in Computing Power Networks

Jingzhao Xie, Zhenglian Li, Gang Sun +2

Computing Power Network (CPN) unifies wide-area computing resources through coordinated network control, while cloud-native abstractions enable flexible resource orchestration and…

cs.DC2025

Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training

Wenjiao Feng, Rongxing Xiao, Zonghang Li +6

Node and link churn in multi-party, cross-region clusters over wide-area networks (WANs) often disrupts distributed training. However, checkpoint-based recovery and cloud-centric a…

cs.NI2025

Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration

Haoxiang Luo, Yinqiu Liu, Ruichen Zhang +7

Edge computing enables real-time data processing closer to its source, thus improving the latency and performance of edge-enabled AI applications. However, traditional AI models of…

cs.DC2024

TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices

Zonghang Li, Wenjiao Feng, Mohsen Guizani +1

Large model inference is shifting from cloud to edge due to concerns about the privacy of user interaction data. However, edge devices often struggle with limited computing power,…