collaborators

5 papers

cs.CV2025

Pre-training Vision Transformers with Formula-driven Supervised Learning

Hirokatsu Kataoka, Sora Takashima, Ryo Hayamizu +6

In the present work, we show that the performance of formula-driven supervised learning (FDSL) can match or even exceed that of ImageNet-21k and can approach that of the JFT-300M d…

cs.SD2025

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data

Yuto Shibata, Keitaro Tanaka, Yoshiaki Bando +3

In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametr…

cs.CV2025

Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering

Erika Mori, Yue Qiu, Hirokatsu Kataoka +1

Social intelligence, the ability to interpret emotions, intentions, and behaviors, is essential for effective communication and adaptive responses. As robots and AI systems become…

cs.CV2025

MoireDB: Formula-generated Interference-fringe Image Dataset

Yuto Matsuo, Ryo Hayamizu, Hirokatsu Kataoka +1

Image recognition models have struggled to treat recognition robustness to real-world degradations. In this context, data augmentation methods like PixMix improve robustness but re…

cs.CV2024

Human Action Recognition without Human

Hirokatsu Kataoka, Kensho Hara, Yutaka Satoh

The objective of this paper is to evaluate "human action recognition without human". Motion representation is frequently discussed in human action recognition. We have examined sev…