papers

Publications (14)

cs.CV2021

Auto Graph Encoder-Decoder for Neural Network Pruning

Sixing Yu, Arya Mazaheri, Ali Jannesari

Model compression aims to deploy deep neural networks (DNN) on mobile devices with limited computing and storage resources. However, most of the existing model compression methods…

cs.LG2024

Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models

Sixing Yu, J. Pablo Muñoz, Ali Jannesari

Foundation Models (FMs), such as LLaMA, BERT, GPT, ViT, and CLIP, have demonstrated remarkable success in a wide range of applications, driven by their ability to leverage vast amo…

cs.LG2025

PerfMamba: Performance Analysis and Pruning of Selective State Space Models

Abdullah Al Asif, Mobina Kashaniyan, Sixing Yu +2

Recent advances in sequence modeling have introduced selective SSMs as promising alternatives to Transformer architectures, offering theoretical computational efficiency and sequen…

cs.DC2022

Towards Seamless Management of AI Models in High-Performance Computing

Sixing Yu, Murali Emani, Chunhua Liao +4

With the increasing prevalence of artificial intelligence (AI) in diverse science/engineering communities, AI models emerge on an unprecedented scale among various domains. However…

cs.CL2024

PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation

Branden Butler, Sixing Yu, Arya Mazaheri +1

Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from C…

cs.LG2023

Heterogeneous Federated Learning using Dynamic Model Pruning and Adaptive Gradient

Sixing Yu, Phuong Nguyen, Ali Anwar +1

Federated Learning (FL) has emerged as a new paradigm for training machine learning models distributively without sacrificing data security and privacy. Learning models on edge dev…