3 papers
cs.CV2025
MINR: Implicit Neural Representations with Masked Image Modelling
Sua Lee, Joonhun Lee, Myungjoo Kang
Self-supervised learning methods like masked autoencoders (MAE) have shown significant promise in learning robust feature representations, particularly in image reconstruction-base…
cs.CV2025
PointT2I: LLM-based text-to-image generation via keypoints
Taekyung Lee, Donggyu Lee, Myungjoo Kang
Text-to-image (T2I) generation model has made significant advancements, resulting in high-quality images aligned with an input prompt. However, despite T2I generation's ability to…
cs.SD2025
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
Suhwan Choi, Kyu Won Kim, Myungjoo Kang
We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions…