4 papers
Human-CLAP: Human-perception-based contrastive language-audio pretraining
Taisei Takano, Yuki Okamoto, Yusuke Kanamori +3
Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, h…
Color-based Emotion Representation for Speech Emotion Recognition
Ryotaro Nagase, Ryoichi Takashima, Yoichi Yamashita
Speech emotion recognition (SER) has traditionally relied on categorical or dimensional labels. However, this technique is limited in representing both the diversity and interpreta…
Construction and Analysis of Impression Caption Dataset for Environmental Sounds
Yuki Okamoto, Ryotaro Nagase, Minami Okamoto +4
Some datasets with the described content and order of occurrence of sounds have been released for conversion between environmental sound and text. However, there are very few texts…
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
Ryotaro Nagase, Takashi Sumiyoshi, Natsuo Yamashita +2
This paper proposes a zero-shot speech emotion recognition (SER) method that estimates emotions not previously defined in the SER model training. Conventional methods are limited t…