5 papers
EgoX: Egocentric Video Generation from a Single Exocentric Video
Taewoong Kang, Kinam Kim, Dohyeon Kim +3
Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (fir…
Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
Daehoon Gwak, Minseo Jung, Junwoo Park +4
Masked diffusion models (MDMs) offer a promising non-autoregressive alternative for large language modeling. Standard decoding methods for MDMs, such as confidence-based sampling,…
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
Sungwon Hwang, Hyojin Jang, Kinam Kim +2
Fine-tuning Video Diffusion Models (VDMs) at the user level to generate videos that reflect specific attributes of training data presents notable challenges, yet remains underexplo…
SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent Representation
Minho Park, Taewoong Kang, Jooyeol Yun +2
The increasing demand for AR/VR applications has highlighted the need for high-quality content, such as 360° live wallpapers. However, generating high-quality 360° panoramic conten…
Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling
Daehoon Gwak, Junwoo Park, Minho Park +4
Predicting future international events from textual information, such as news articles, has tremendous potential for applications in global policy, strategic decision-making, and g…