activity
20232026
most citedAdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference

18 citations · 19 across the 23 of their papers we have counts for

collaborators

17 papers

cs.AR2026

Aging Aware Adaptive Voltage Scaling for Reliable and Efficient AI Accelerators

Tong Xie, Zuodong Zhang, Chao Yang +3

Deep neural networks (DNNs) have showcased remarkable performance across various tasks and are widely deployed on AI accelerators fabricated in advanced technology nodes for effici…

cs.AR2026

DRIFT: Harnessing Inherent Fault Tolerance for Efficient and Reliable Diffusion Model Inference

Jinqi Wen, Tong Xie, Runsheng Wang +1

Diffusion model deployment has been suffering from high energy consumption and inference latency despite its superior performance in visual generation tasks. Dynamic voltage and fr…

cs.AR2026

The Quest for Reliable AI Accelerators: Cross-Layer Evaluation and Design Optimization

Meng Li, Tong Xie, Zuodong Zhang +1

As the CMOS technology pushes to the nanoscale, aging effects and process variations have become increasingly pronounced, posing significant reliability challenges for AI accelerat…

cs.AR2026

CREATE: Cross-Layer Resilience Characterization and Optimization for Efficient yet Reliable Embodied AI Systems

Tong Xie, Yijiahao Qi, Jinqi Wen +9

Embodied Artificial Intelligence (AI) has recently attracted significant attention as it bridges AI with the physical world. Modern embodied AI systems often combine a Large Langua…

cs.PF2025

HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing

Haochen Huang, Shuzhang Zhong, Zhe Zhang +5

Large Language Models (LLMs) with Mixture-of-Expert (MoE) architectures achieve superior model performance with reduced computation costs, but at the cost of high memory capacity a…

cs.CR2025

Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPC

Tianshi Xu, Wen-jie Lu, Jiangrui Yu +4

This paper presents an efficient framework for private Transformer inference that combines Homomorphic Encryption (HE) and Secure Multi-party Computation (MPC) to protect data priv…