Showing cs.ARShow all
2 papers · 1 filter
cs.AR2025
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
Chun-Ting Chen, HanGyeol Mun, Jian Meng +2
Edge inference for large language models (LLM) offers secure, low-latency, and cost-effective inference solutions. We emphasize that an edge accelerator should achieve high area ef…
cs.AR2024
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
Jian Meng, Yuan Liao, Anupreetham Anupreetham +5
The development of model compression is continuously motivated by the evolution of various neural network accelerators with ASIC or FPGA. On the algorithm side, the ultimate goal o…