1 citations · 1 across the 1 of their papers we have counts for
4 papers
Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun +1
This paper introduces a cutting-edge approach to cross-modal interaction for tiny object detection by combining semantic-guided natural language processing with advanced visual rec…
FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
This paper introduces FaceLiVT, a lightweight yet powerful face recognition model that integrates a hybrid Convolution Neural Network (CNN)-Transformer architecture with an innovat…
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving
Novendra Setyawan, Ghufron Wahyu Kurniawan, Chi-Chia Sun +2
The perception system is a a critical role of an autonomous driving system for ensuring safety. The driving scene perception system fundamentally represents an object detection tas…
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
The Vision Transformer (ViT) has demonstrated state-of-the-art performance in various computer vision tasks, but its high computational demands make it impractical for edge devices…