From the 2 of 9 linked papers with an AI index.
9 papers
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
Ku Onoda, Paavo Parmas, Hiroki Furuta +4
Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. Th…
Visual Access Boundaries in Vision-Language Model Reasoning
Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi +3
The paper investigates how chain-of-thought prompting works in vision-language models by introducing a visual access sweep that masks attention to image tokens, defining a visual a…
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki +1
Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in extending these models to video generation.…
CLIP-like Model as a Foundational Density Ratio Estimator
Fumiya Uchiyama, Rintaro Yanagi, Shohei Taniguchi +5
Density ratio estimation is a core concept in statistical machine learning because it provides a unified mechanism for tasks such as importance weighting, divergence estimation, an…
Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability
Bum Jun Kim, Shohei Taniguchi, Makoto Kawano +2
Training divergence in transformers wastes compute, yet practitioners discover instability only after expensive runs begin. They therefore need an expected probability of failure f…
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers
Atsushi Shimizu, Shohei Taniguchi, Yutaka Matsuo
Length generalization is the ability of language models to maintain performance on inputs longer than those seen during pretraining. In this work, we introduce a simple yet powerfu…