audio watermarking 1diffusion models 1latent space embedding 1music generation 1robustness to attacks 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.SD2026
MusicMark: A Robust Generative Watermarking Framework for Music Generation
Seohwan Yun, Jeeyoung Yun, Yongjin Kim +2
The paper introduces MusicMark, a framework that embeds watermarks directly into the latent space of diffusion-based music generation models, making the watermarks robust to transf…
cs.CL2026
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
Daejin Jo, Jeeyoung Yun, Byungseok Roh +1
With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modaliti…
cs.CV2026
FiRe: Fine-grained Multimodal Reasoning for Enhanced Image Generation
Yongjin Kim, Yoonjin Oh, Yerin Kim +5
With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced significantly. However, des…