1 citations · 1 across the 4 of their papers we have counts for
14 papers
DeltaServe: Host-Agnostic Co-Serving of Inference and Fine-Tuning for LLMs
Jiaxuan Chen, Jianshu She, Ye Yuan +5
LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. We present DeltaServe, a h…
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
Zonghang Li, Tao Li, Wenjiao Feng +8
On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low throughput and capability. To overcome th…
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
Akhmed Sakip, Erland Hilman Fuadi, Omar Sayedelahl +6
Training large language models requires jointly configuring two interdependent aspects of the system: the global batch size, which governs statistical efficiency, and the 3D parall…
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
Chong Tian, Yu Wang, Chenxu Yang +5
Short-form video platforms are major channels for news but also fertile ground for multimodal misinformation where each modality appears plausible alone yet cross-modal relationshi…
LAPS: A Length-Aware-Prefill LLM Serving System
Jianshu She, Zonghang Li, Hongchao Du +7
LAPS identifies and disaggregates requests with different prompt lengths in LLM serving to reduce TTFT latency. While recent systems have decoupled the prefill and decode stages to…
Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
Wenjiao Feng, Rongxing Xiao, Zonghang Li +6
Node and link churn in multi-party, cross-region clusters over wide-area networks (WANs) often disrupts distributed training. However, checkpoint-based recovery and cloud-centric a…