4 papers
DMoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
Haodong Wang, Qihua Zhou, Zicong Hong +1
The mixture of experts (MoE) model is a sparse variant of large language models (LLMs), designed to hold a better balance between intelligent capability and computational overhead.…
Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
Bohai Gu, Hao Luo, Song Guo +2
The text-guided video inpainting technique has significantly improved the performance of content generation applications. A recent family for these improvements uses diffusion mode…
The MoE-Empowered Edge LLMs Deployment: Architecture, Challenges, and Opportunities
Ning Li, Song Guo, Tuo Zhang +5
The powerfulness of LLMs indicates that deploying various LLMs with different scales and architectures on end, edge, and cloud to satisfy different requirements and adaptive hetero…
Mjolnir: Breaking the Shield of Perturbation-Protected Gradients via Adaptive Diffusion
Xuan Liu, Siqi Cai, Qihua Zhou +3
Perturbation-based mechanisms, such as differential privacy, mitigate gradient leakage attacks by introducing noise into the gradients, thereby preventing attackers from reconstruc…