Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
Realizable N:M Sparse Transformer Inference via Search-Kernel Co-Design
Yiming Liu, Wenqi Lou, Zhiguang Wang +4
Vision Transformers (ViTs) achieve strong accuracy but incur high inference latency. Semi-structured N:M sparsity can reduce arithmetic cost, yet its theoretical savings often fail…
cs.AR2025
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
Jiale Dong, Hao Wu, Zihao Wang +5
Vision Transformers (ViTs) exhibit superior performance in computer vision tasks but face deployment challenges on resource-constrained devices due to high computational/memory dem…
cs.AR2025
UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGA
Jiale Dong, Wenqi Lou, Zhendong Zheng +4
Compared to traditional Vision Transformers (ViT), Mixture-of-Experts Vision Transformers (MoE-ViT) are introduced to scale model size without a proportional increase in computatio…