Publications (14)
Auto Graph Encoder-Decoder for Neural Network Pruning
Sixing Yu, Arya Mazaheri, Ali Jannesari
Model compression aims to deploy deep neural networks (DNN) on mobile devices with limited computing and storage resources. However, most of the existing model compression methods…
Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models
Sixing Yu, J. Pablo Muñoz, Ali Jannesari
Foundation Models (FMs), such as LLaMA, BERT, GPT, ViT, and CLIP, have demonstrated remarkable success in a wide range of applications, driven by their ability to leverage vast amo…
PerfMamba: Performance Analysis and Pruning of Selective State Space Models
Abdullah Al Asif, Mobina Kashaniyan, Sixing Yu +2
Recent advances in sequence modeling have introduced selective SSMs as promising alternatives to Transformer architectures, offering theoretical computational efficiency and sequen…
Towards Seamless Management of AI Models in High-Performance Computing
Sixing Yu, Murali Emani, Chunhua Liao +4
With the increasing prevalence of artificial intelligence (AI) in diverse science/engineering communities, AI models emerge on an unprecedented scale among various domains. However…
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
Branden Butler, Sixing Yu, Arya Mazaheri +1
Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from C…
Heterogeneous Federated Learning using Dynamic Model Pruning and Adaptive Gradient
Sixing Yu, Phuong Nguyen, Ali Anwar +1
Federated Learning (FL) has emerged as a new paradigm for training machine learning models distributively without sacrificing data security and privacy. Learning models on edge dev…