5 papers
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
Takafumi Moriya, Masato Mimura, Tomohiro Tanaka +3
This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist tempor…
Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
Satoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda +7
Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image…
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
Ryo Masumura, Shota Orihashi, Mana Ihori +6
This paper proposes a joint modeling method of the Big Five, which has long been studied, and HEXACO, which has recently attracted attention in psychology, for automatically recogn…
Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
Mana Ihori, Taiga Yamane, Naotaka Kawata +5
This paper proposes a personalization method for speech emotion recognition (SER) through in-context learning (ICL). Since the expression of emotions varies from person to person,…
MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
Taiga Yamane, Satoshi Suzuki, Ryo Masumura +5
Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view (BEV) from multi-view images. In MVPD, end-to-end trainable deep learning methods…