43 papers
Distilling Physical Priors into Streaming World Models
Liangliang Zhao, Junying Wang, Danni Yang +5
Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical c…
HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding
Wenhui Liao, Hongliang Li, Pengyu Xie +15
Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…
Beyond the LUMIR challenge: The pathway to foundational registration models
Junyu Chen, Shuwen Wei, Joel Honkamaa +33
Medical image challenges have played a transformative role in advancing the field, catalyzing innovation and establishing new performance benchmarks. Image registration, a foundati…
Large-Scale Deployment and Analytical Implications of Structured Quality Control in Diffusion Magnetic Resonance Imaging
Michael E. Kim, Chenyu Gao, Karthik Ramadass +17
Purpose: Diffusion MRI (dMRI) provides a diverse set of quantitative measures and derived datatypes to assess white matter microstructure and macrostructure. Coupled with the incre…
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
Siqi Luo, Jianghan Shen, Yi Xin +9
Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One prima…
StableI2I: Spotting Unintended Changes in Image-to-Image Transition
Jiayang Li, Shuo Cao, Xiaohui Li +6
In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. H…