2 papers
cs.AR2026
SATA: Sparsity-Aware Scheduling for Selective Token Attention
Zhenkun Fan, Zishen Wan, Che-Kai Liu +5
Transformers have become the foundation of numerous state-of-the-art AI models across diverse domains, thanks to their powerful attention mechanism for modeling long-range dependen…
cs.CV2024
Neural Architecture Search of Hybrid Models for NPU-CIM Heterogeneous AR/VR Devices
Yiwei Zhao, Ziyun Li, Win-San Khwa +10
Low-Latency and Low-Power Edge AI is essential for Virtual Reality and Augmented Reality applications. Recent advances show that hybrid models, combining convolution layers (CNN) a…