Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration
Yuan Liao, Jae-sun Seo
The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerato…
cs.AR2025
Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
Chun-Ting Chen, HanGyeol Mun, Jian Meng +2
Edge inference for large language models (LLM) offers secure, low-latency, and cost-effective inference solutions. We emphasize that an edge accelerator should achieve high area ef…
cs.AR2024
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design
Jian Meng, Yuan Liao, Anupreetham Anupreetham +5
The development of model compression is continuously motivated by the evolution of various neural network accelerators with ASIC or FPGA. On the algorithm side, the ultimate goal o…