7 papers
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
Ku Onoda, Paavo Parmas, Hiroki Furuta +4
Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. Th…
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki +1
Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in extending these models to video generation.…
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
Yuta Oshima, Daiki Miyake, Kohsei Matsutani +4
Recent text-to-image generation models have acquired the ability of multi-reference generation and editing; that is, to inherit the appearance of subjects from multiple reference i…
WorldPack: Dynamic Frame Compression for Long-context Video World Modeling
Yuta Oshima, Yusuke Iwasawa, Masahiro Suzuki +2
Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation action…
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo +1
The remarkable progress in text-to-video diffusion models enables the generation of photorealistic videos, although the content of these generated videos often includes unnatural m…
ADOPT: Modified Adam Can Converge with Any with the Optimal Rate
Shohei Taniguchi, Keno Harada, Gouki Minegishi +7
Adam is one of the most popular optimization algorithms in deep learning. However, it is known that Adam does not converge in theory unless choosing a hyperparameter, i.e., ,…