3 papers
cs.CV2025
Action tube generation by person query matching for spatio-temporal action detection
Kazuki Omi, Jion Oshima, Toru Tamaki
This paper proposes a method for spatio-temporal action detection (STAD) that directly generates action tubes from the original video without relying on post-processing steps such…
cs.CV2024
Shift and matching queries for video semantic segmentation
Tsubasa Mizuno, Toru Tamaki
Video segmentation is a popular task, but applying image segmentation models frame-by-frame to videos does not preserve temporal consistency. In this paper, we propose a method to…
cs.CV2024
Query matching for spatio-temporal action detection with query-based object detector
Shimon Hori, Kazuki Omi, Toru Tamaki
In this paper, we propose a method that extends the query-based object detection model, DETR, to spatio-temporal action detection, which requires maintaining temporal consistency i…