collaborators

6 papers

cs.CV2026

PowerCLIP: Powerset Alignment for Contrastive Pre-Training

Masaki Kawamura, Nakamasa Inoue, Rintaro Yanagi +2

Contrastive vision-language pre-training frameworks such as CLIP have demonstrated impressive zero-shot performance across a range of vision-language tasks. Recent studies have sho…

cs.SD2026

AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference

Risa Shinoda, Kaede Shiohara, Nakamasa Inoue +2

Animal vocalizations provide crucial insights for wildlife assessment, particularly in complex environments such as forests, aiding species identification and ecological monitoring…

cs.CV2025

Pre-training Vision Transformers with Formula-driven Supervised Learning

Hirokatsu Kataoka, Sora Takashima, Ryo Hayamizu +6

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k and can approach that of the JFT-300M d…

cs.CV2025

AgroBench: Vision-Language Model Benchmark in Agriculture

Risa Shinoda, Nakamasa Inoue, Hirokatsu Kataoka +2

Precise automated understanding of agricultural tasks such as disease identification is essential for sustainable crop production. Recent advances in vision-language models (VLMs)…

cs.CV2025

AnimalClue: Recognizing Animals by their Traces

Risa Shinoda, Nakamasa Inoue, Iro Laina +2

Wildlife observation plays an important role in biodiversity conservation, necessitating robust methodologies for monitoring wildlife populations and interspecies interactions. Rec…

cs.CV2025

On the Relationship Between Double Descent of CNNs and Shape/Texture Bias Under Learning Process

Shun Iwase, Shuya Takahashi, Nakamasa Inoue +3

The double descent phenomenon, which deviates from the traditional bias-variance trade-off theory, attracts considerable research attention; however, the mechanism of its occurrenc…