6 papers
HyQuant: Hybrid-Precision Quantization for LLM Attention
Jiatong Ding, Bingxin Xing, Yu Zhang +9
Quantization has been widely adopted in LLM training and inference to reduce cost and improve efficiency. However, low-bit quantization of the \emph{attention} module often introdu…
Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving
Weinan Liu, Zeyuan Ding, Dian Ding +5
Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained oppo…
SWIFT: Spatio-temporal Wavelet Integrated Forecasting Framework for Workload Traces
Zeyuan Ding, Lingfeng Zheng, Dian Ding +1
Accurate cloud workload forecasting is pivotal for efficient resource management but remains challenging as workloads are highly volatile and prone to sudden bursts. Although wavel…
Wideband RF Radiance Field Modeling Using Frequency-embedded 3D Gaussian Splatting
Zechen Li, Lanqing Yang, Yiheng Bian +6
Indoor environments typically contain diverse RF signals distributed across multiple frequency bands, including NB-IoT, Wi-Fi, and millimeter-wave. Consequently, wideband RF modeli…
Push the Limit of Multi-modal Emotion Recognition by Prompting LLMs with Receptive-Field-Aware Attention Weighting
Han Zhang, Yu Lu, Liyun Zhang +5
Understanding the emotions in a dialogue usually requires external knowledge to accurately understand the contents. As the LLMs become more and more powerful, we do not want to set…
VCEMO: Multi-Modal Emotion Recognition for Chinese Voiceprints
Jinghua Tang, Liyun Zhang, Yu Lu +6
Emotion recognition can enhance humanized machine responses to user commands, while voiceprint-based perception systems can be easily integrated into commonly used devices like sma…