2 citations · 8 across the 26 of their papers we have counts for
6 papers · 1 filter
MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model
Taiga Yamane, Satoshi Suzuki, Ryo Masumura +4
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adopt a unified framework that pr…
Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda +7
Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image…
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
Ryo Masumura, Shota Orihashi, Mana Ihori +6
This paper proposes a joint modeling method of the Big Five, which has long been studied, and HEXACO, which has recently attracted attention in psychology, for automatically recogn…
MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
Taiga Yamane, Satoshi Suzuki, Ryo Masumura +5
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view (BEV) from multi-view images. In MVPD, end-to-end trainable deep learning methods…
Ladder Siamese Network: a Method and Insights for Multi-level Self-Supervised Learning
Ryota Yoshihashi, Shuhei Nishimura, Dai Yonebayashi +3
Siamese-network-based self-supervised learning (SSL) suffers from slow convergence and instability in training. To alleviate this, we propose a framework to exploit intermediate se…
Context-Free TextSpotter for Real-Time and Mobile End-to-End Text Detection and Recognition
Ryota Yoshihashi, Tomohiro Tanaka, Kenji Doi +2
In the deployment of scene-text spotting systems on mobile platforms, lightweight models with low computation are preferable. In concept, end-to-end (E2E) text spotting is suitable…