4 papers
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
Afsara Benazir, Felix Xiaozhu Lin
Apple Neural Engine (ANE) is a dedicated neural processing unit (NPU) present in every Apple Silicon chip. Mixture-of-Experts (MoE) LLMs improve inference efficiency via sparse act…
Safeguarding Privacy in Edge Speech Understanding with Tiny Foundation Models
Afsara Benazir, Felix Xiaozhu Lin
Robust speech recognition systems rely on cloud service providers for inference. It needs to ensure that an untrustworthy provider cannot deduce the sensitive content in speech. Sa…
Profiling Large Language Model Inference on Apple Silicon: A Quantization Perspective
Afsara Benazir, Felix Xiaozhu Lin
A systematic understanding of Apple Silicon is lacking in the current landscape of hardware efficiency; research focus is largely centered on accelerating GPUs for large-scale trai…
SelectFormer: Private and Practical Data Selection for Transformers
Xu Ouyang, Felix Xiaozhu Lin, Yangfeng Ji
Critical to a free data market is , i.e. the model owner selects and then appraises training data from the data owner before both parties commit to…