21 citations · 25 across the 2 of their papers we have counts for
4 papers
Compass: A Decentralized Scheduler for Latency-Sensitive ML Workflows
Yuting Yang, Andrea Merlina, Weijia Song +3
We consider ML query processing in distributed systems where GPU-enabled workers coordinate to execute complex queries: a computing style often seen in applications that interact w…
Towards Efficient Verification of Quantized Neural Networks
Pei Huang, Haoze Wu, Yuting Yang +4
Quantization replaces floating point arithmetic with integer arithmetic in deep neural network models, providing more efficient on-device inference with less power and memory. In t…
Enhancing the Unified Streaming and Non-streaming Model with Contrastive Learning
Yuting Yang, Yuke Li, Binbin Du
The unified streaming and non-streaming speech recognition model has achieved great success due to its comprehensive capabilities. In this paper, we propose to improve the accuracy…
FreConv: Frequency Branch-and-Integration Convolutional Networks
Zhaowen Li, Xu Zhao, Peigeng Ding +4
Recent researches indicate that utilizing the frequency information of input data can enhance the performance of networks. However, the existing popular convolutional structure is…