works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.CV2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +3

Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelope…

cs.CV2026

Confidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics

Yi Tang Soon, Jun-Wei Hsieh

The paper investigates why confidence scores from open‑vocabulary object detectors built on CLIP are biased by object size and query specificity, and proposes a simple temperature‑…

cs.CV2026

TinyFormer: Preserving Tiny Objects in YOLO-DETR Hybrid Real-time Detectors

Jun-Wei Hsieh, Meng-Yu Kao, Ghufron Wahyu Kurniawan +1

YOLO-series and DETR-based detectors struggle with tiny-object detection. YOLO-style models benefit from efficient dense prediction, but their large-stride backbones may suppress t…

cs.CV2026

MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2

The Vision Transformer (ViT) achieves remarkable accuracy across visual tasks but remains computationally expensive for edge deployment. This paper presents MicroViTv2, a lightweig…

cs.CV2026

FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu +2

Lightweight face recognition is increasingly important for deployment on edge and mobile devices, where strict constraints on latency, memory, and energy consumption must be met al…

cs.CV2025

Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection

Xian-Hong Huang, Hui-Kai Su, Chi-Chia Sun +1

This paper introduces a cutting-edge approach to cross-modal interaction for tiny object detection by combining semantic-guided natural language processing with advanced visual rec…