10 citations · 15 across the 5 of their papers we have counts for
5 papers
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
Jaehyeon Kim, Keon Lee, Seungjun Chung +1
With the emergence of neural audio codecs, which encode multiple streams of discrete tokens from audio, large language models have recently gained attention as a promising approach…
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
Jongho Park, Jaeseung Park, Zheyang Xiong +5
State-space models (SSMs), such as Mamba (Gu & Dao, 2023), have been proposed as alternatives to Transformer networks in language modeling, by incorporating gating, convolutions, a…
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
Inkyu Park, Jaewoong Cho
Speech-driven 3D facial animation is challenging due to the scarcity of large-scale visual-audio datasets despite extensive research. Most prior works, typically focused on learnin…
Addressing Feature Imbalance in Sound Source Separation
Jaechang Kim, Jeongyeon Hwang, Soheun Yi +2
Neural networks often suffer from a feature preference problem, where they tend to overly rely on specific features to solve a task while disregarding other features, even if those…
Mini-Batch Optimization of Contrastive Loss
Jaewoong Cho, Kartik Sreenivasan, Keon Lee +7
Contrastive learning has gained significant attention as a method for self-supervised learning. The contrastive loss function ensures that embeddings of positive sample pairs (e.g.…