11 citations · 11 across the 2 of their papers we have counts for
3 papers
eess.IV2026★ 11 cited
A 28nm 0.22μJ/token memory-compute-intensity-aware CNN-Transformer accelerator with hybrid-attention-based layer-fusion and cascaded pruning for semantic-segmentation
Pingcheng Dong, Yonghao Tan, Xuejiao Liu +14
This work presents a 28nm 13.93mm2 CNN-Transformer accelerator for semantic segmentation, achieving 3.86-to-10.91x energy reduction over previous designs. It features a hybrid atte…
cs.AR2025
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
Kunming Shao, Zhipeng Liao, Jiangnan Yu +9
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, ene…
cs.AR2025
A Flexible Precision Scaling Deep Neural Network Accelerator with Efficient Weight Combination
Liang Zhao, Kunming Shao, Fengshi Tian +3
Deploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is…