From the 1 of 2 linked papers with an AI index.
4 papers
DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
Takyoung Kim, Kang-wook Kim, Sang Hoon Woo +3
The paper presents DuplexGen, a framework that uses a small set of human preference annotations to adaptively generate turn‑taking behaviors in AI‑human dialogues across different…
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
Jaeyeon Kim, Heeseung Yun, Sang Hoon Woo +2
Large audio language models (LALMs) extend language understanding into the auditory domain, yet their ability to perform low-level listening, such as pitch and duration detection,…
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance
Jaeyeon Kim, Minjeon Jeon, Jaeyoon Jung +2
In this work, we aim to analyze and optimize the EnCLAP framework, a state-of-the-art model in automated audio captioning. We investigate the impact of modifying the acoustic encod…
Expanding on EnCLAP with Auxiliary Retrieval Model for Automated Audio Captioning
Jaeyeon Kim, Jaeyoon Jung, Minjeong Jeon +2
In this technical report, we describe our submission to DCASE2024 Challenge Task6 (Automated Audio Captioning) and Task8 (Language-based Audio Retrieval). We develop our approach b…