From the 1 of 9 linked papers with an AI index.
9 papers
CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +3
Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelope…
Confidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics
Yi Tang Soon, Jun-Wei Hsieh
The paper investigates why confidence scores from open‑vocabulary object detectors built on CLIP are biased by object size and query specificity, and proposes a simple temperature‑…
TinyFormer: Preserving Tiny Objects in YOLO-DETR Hybrid Real-time Detectors
Jun-Wei Hsieh, Meng-Yu Kao, Ghufron Wahyu Kurniawan +1
YOLO-series and DETR-based detectors struggle with tiny-object detection. YOLO-style models benefit from efficient dense prediction, but their large-stride backbones may suppress t…
MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
The Vision Transformer (ViT) achieves remarkable accuracy across visual tasks but remains computationally expensive for edge deployment. This paper presents MicroViTv2, a lightweig…
FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition
Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2
Lightweight face recognition is increasingly important for deployment on edge and mobile devices, where strict constraints on latency, memory, and energy consumption must be met al…
Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun +1
This paper introduces a cutting-edge approach to cross-modal interaction for tiny object detection by combining semantic-guided natural language processing with advanced visual rec…