activity
20192025
most citedRWCP-SSD-Onomatopoeia: Onomatopoeic Word Dataset for Environmental Sound Synthesis

2 citations · 3 across the 9 of their papers we have counts for

collaborators
Showing cs.SDShow all

11 papers · 1 filter

cs.SD2024

Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation

Junwon Lee, Modan Tailleur, Laurie M. Heller +5

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene…

cs.SD2024

Construction and Analysis of Impression Caption Dataset for Environmental Sounds

Yuki Okamoto, Ryotaro Nagase, Minami Okamoto +4

Some datasets with the described content and order of occurrence of sounds have been released for conversion between environmental sound and text. However, there are very few texts…

cs.SD2024

Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant

Modan Tailleur, Junwon Lee, Mathieu Lagrange +4

This paper explores whether considering alternative domain-specific embeddings to calculate the Fréchet Audio Distance (FAD) metric can help the FAD to correlate better with percep…

cs.SD2023

CAPTDURE: Captioned Sound Dataset of Single Sources

Yuki Okamoto, Kanta Shimonishi, Keisuke Imoto +3

In conventional studies on environmental sound separation and synthesis using captions, datasets consisting of multiple-source sounds with their captions were used for model traini…

cs.SD2023

Environmental sound synthesis from vocal imitations and sound event labels

Yuki Okamoto, Keisuke Imoto, Shinnosuke Takamichi +3

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effect…

cs.SD2023

Foley Sound Synthesis at the DCASE 2023 Challenge

Keunwoo Choi, Jaekwon Im, Laurie Heller +5

The addition of Foley sound effects during post-production is a common technique used to enhance the perceived acoustic properties of multimedia content. Traditionally, Foley sound…