From the 1 of 8 linked papers with an AI index.
8 papers
Visual Access Boundaries in Vision-Language Model Reasoning
Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi +3
The paper investigates how chain-of-thought prompting works in vision-language models by introducing a visual access sweep that masks attention to image tokens, defining a visual a…
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
Tomoshi Iiyama, Masahiro Suzuki, Yutaka Matsuo
Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks. However, their performance hinges on how…
CLIP-like Model as a Foundational Density Ratio Estimator
Fumiya Uchiyama, Rintaro Yanagi, Shohei Taniguchi +5
Density ratio estimation is a core concept in statistical machine learning because it provides a unified mechanism for tasks such as importance weighting, divergence estimation, an…
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
Yuta Oshima, Daiki Miyake, Kohsei Matsutani +4
Recent text-to-image generation models have acquired the ability of multi-reference generation and editing; that is, to inherit the appearance of subjects from multiple reference i…
WorldPack: Dynamic Frame Compression for Long-context Video World Modeling
Yuta Oshima, Yusuke Iwasawa, Masahiro Suzuki +2
Video world models have attracted significant attention for their ability to produce high-fidelity future visual observations conditioned on past observations and navigation action…
When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks
Stefano Ferraro, Akihiro Nakano, Masahiro Suzuki +1
Object-centric world models (OCWM) aim to decompose visual scenes into object-level representations, providing structured abstractions that could improve compositional generalizati…