3 papers
cs.LG2026
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs
Yu Luo, Bo Dong, Wenhua Cheng +1
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint, and throughput. However, conventional qua…
cs.CV2024
DEGAS: Detailed Expressions on Full-Body Gaussian Avatars
Zhijing Shao, Duotun Wang, Qing-Yao Tian +7
Although neural rendering has made significant advances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remai…
cs.LG2023
Efficient LLM Inference on CPUs
Haihao Shen, Hanwen Chang, Bo Dong +2
Large language models (LLMs) have demonstrated remarkable performance and tremendous potential across a wide range of tasks. However, deploying these models has been challenging du…