activity
20192026
most citedMVA2023 Small Object Detection Challenge for Spotting Birds: Dataset, Methods, and Results

20 citations · 79 across the 14 of their papers we have counts for

collaborators

15 papers

cs.CV2026

ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation

Trung Thanh Nguyen, Tuan-Anh Vu, Duc Viet Le +4

Semantic and instance segmentation of terrestrial and drone LiDAR point clouds is emerging as a transformative approach for converting the complex 3D structure of forests into acti…

cs.CV2025★ 11 cited

Small Object Detection for Birds with Swin Transformer

Da Huo, Marc A. Kastner, Tingwei Liu +4

Object detection is the task of detecting objects in an image. In this task, the detection of small objects is particularly difficult. Other than the small size, it is also accompa…

cs.CV2024★ 9 cited

Action Selection Learning for Multi-label Multi-view Action Recognition

Trung Thanh Nguyen, Yasutomo Kawanishi, Takahiro Komamizu +1

Multi-label multi-view action recognition aims to recognize multiple concurrent or sequential actions from untrimmed videos captured by multiple cameras. Existing work has focused…

cs.CV2024★ 2 cited

Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association

Tingwei Liu, Yasutomo Kawanishi, Takahiro Komamizu +1

This paper focuses on tracking birds that appear small in a panoramic video. When the size of the tracked object is small in the image (small object tracking) and move quickly, obj…

cs.CV2024

One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features

Trung Thanh Nguyen, Yasutomo Kawanishi, Takahiro Komamizu +1

Open-vocabulary Temporal Action Detection (Open-vocab TAD) is an advanced video analysis approach that expands Closed-vocabulary Temporal Action Detection (Closed-vocab TAD) capabi…

cs.CV2024

Multi-View Video-Based Learning: Leveraging Weak Labels for Frame-Level Perception

Vijay John, Yasutomo Kawanishi

For training a video-based action recognition model that accepts multi-view video, annotating frame-level labels is tedious and difficult. However, it is relatively easy to annotat…