8 papers
CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +3
Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelope…
MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
The Vision Transformer (ViT) achieves remarkable accuracy across visual tasks but remains computationally expensive for edge deployment. This paper presents MicroViTv2, a lightweig…
FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
Lightweight face recognition is increasingly important for deployment on edge and mobile devices, where strict constraints on latency, memory, and energy consumption must be met al…
Fast Person Detection Using YOLOX With AI Accelerator For Train Station Safety
Mas Nurul Achmadiah, Novendra Setyawan, Achmad Arif Bryantono +2
Recently, Image processing has advanced Faster and applied in many fields, including health, industry, and transportation. In the transportation sector, object detection is widely…
RepSFNet : A Single Fusion Network with Structural Reparameterization for Crowd Counting
Mas Nurul Achmadiah, Chi-Chia Sun, Wen-Kai Kuo +1
Crowd counting remains challenging in variable-density scenes due to scale variations, occlusions, and the high computational cost of existing models. To address these issues, we p…
FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
This paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovat…