6 citations · 11 across the 7 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025★ 2 cited
APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM Acceleration
Shaobo Ma, Chao Fang, Haikuo Shao +1
Large language models (LLMs) have revolutionized AI applications, yet their enormous computational demands severely limit deployment and real-time performance. Quantization methods…
cs.LG2024★ 3 cited
Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor Cores
Shaobo Ma, Chao Fang, Haikuo Shao +1
Large language models (LLMs) have been widely applied but face challenges in efficient inference. While quantization methods reduce computational demands, ultra-low bit quantizatio…
cs.LG2024★ 6 cited
Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge Deployment
Yuhao Ji, Chao Fang, Shaobo Ma +2
Transformer models have revolutionized AI tasks, but their large size hinders real-world deployment on resource-constrained and latency-critical edge devices. While binarized Trans…