3 papers
cs.SD2022
Visual onoma-to-wave: environmental sound synthesis from visual onomatopoeias and sound-source images
Hien Ohnaka, Shinnosuke Takamichi, Keisuke Imoto +3
We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e.…
cs.HC2021
HumanACGAN: conditional generative adversarial network with human-based auxiliary classifier and its evaluation in phoneme perception
Yota Ueda, Kazuki Fujii, Yuki Saito +3
We propose a conditional generative adversarial network (GAN) incorporating humans' perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a rea…
cs.SD2019
HumanGAN: generative adversarial network with human-based discriminator and its evaluation in speech perception modeling
Kazuki Fujii, Yuki Saito, Shinnosuke Takamichi +2
We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribut…