activity
20242026
collaborators

6 papers

cs.LG2026

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

Ku Onoda, Paavo Parmas, Hiroki Furuta +4

Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. Th…

cs.LG2026

Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

Hiroki Furuta, Heiga Zen, Dale Schuurmans +4

Large text-to-video models hold immense potential for a wide range of downstream applications. However, they struggle to accurately depict dynamic object interactions, often result…

cs.CV2026

MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation

Yuta Oshima, Daiki Miyake, Kohsei Matsutani +4

Recent text-to-image generation models have acquired the ability of multi-reference generation and editing; that is, to inherit the appearance of subjects from multiple reference i…

cs.CV2025

WorldPack: Dynamic Frame Compression for Long-context Video World Modeling

Yuta Oshima, Yusuke Iwasawa, Masahiro Suzuki +2

Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation action…

cs.CV2025

Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search

Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo +1

The remarkable progress in text-to-video diffusion models enables the generation of photorealistic videos, although the content of these generated videos often includes unnatural m…

cs.LG2024

Geometric-Averaged Preference Optimization for Soft Preference Labels

Hiroki Furuta, Kuang-Huei Lee, Shixiang Shane Gu +4

Many algorithms for aligning LLMs with human preferences assume that human preferences are binary and deterministic. However, human preferences can vary across individuals, and the…