36 papers
Simplex Relaxation for Discrete Diffusion
Jinya Sakurai, Patrick Pynadath, Satoshi Hayakawa +4
Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem…
Are Video Reasoning Models Ready to Go Outside?
Yangfan He, Changgyu Boo, Jaehong Yoon
In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such conditions, their understanding and reasonin…
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
Yuxuan Fan, Gyusik Seo, Jing Hao +3
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinatio…
Confidence-Aware Tool Orchestration for Robust Video Understanding
Yangfan He, Yujin Choi, Jaehong Yoon
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Problem: under realistic perturbations such…
Safe Few-Step Generation via Velocity Editing
Yujin Choi, Jaehong Yoon
Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a small number of sampling steps.…
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
Shoubin Yu, Yue Zhang, Zun Wang +4
Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under unseen or alternative viewpoints. Recent…