From the 1 of 5 linked papers with an AI index.
5 papers
MusicMark: A Robust Generative Watermarking Framework for Music Generation
Seohwan Yun, Jeeyoung Yun, Yongjin Kim +2
The paper introduces MusicMark, a framework that embeds watermarks directly into the latent space of diffusion-based music generation models, making the watermarks robust to transf…
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
Daejin Jo, Jeeyoung Yun, Byungseok Roh +1
With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modaliti…
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
Donghwan Chi, Hyomin Kim, Yoonjin Oh +7
Recently, multimodal large language models (MLLMs) have emerged as a key approach in achieving artificial general intelligence. In particular, vision-language MLLMs have been devel…
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
Yoonjin Oh, Yongjin Kim, Hyomin Kim +2
Recent advances in Multimodal Large Language Models (MLLMs) have enabled unified multimodal understanding and generation. However, they still struggle with fine-grained text-image…
SGPO: Self-Generated Preference Optimization based on Self-Improver
Hyeonji Lee, Daejin Jo, Seohwan Yun +1
Large language models (LLMs), despite their extensive pretraining on diverse datasets, require effective alignment to human preferences for practical and reliable deployment. Conve…