3 papers
cs.CV2026
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki +1
Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in extending these models to video generation.…
cs.CV2025
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo +1
The remarkable progress in text-to-video diffusion models enables the generation of photorealistic videos, although the content of these generated videos often includes unnatural m…
cs.LG2024
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo
Multimodal variational autoencoders (VAEs) aim to capture shared latent representations by integrating information from different data modalities. A significant challenge is accura…