collaborators

43 papers

cs.CV2026

Distilling Physical Priors into Streaming World Models

Liangliang Zhao, Junying Wang, Danni Yang +5

Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical c…

cs.CV2026

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

Wenhui Liao, Hongliang Li, Pengyu Xie +15

Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…

eess.IV2026

Beyond the LUMIR challenge: The pathway to foundational registration models

Junyu Chen, Shuwen Wei, Joel Honkamaa +33

Medical image challenges have played a transformative role in advancing the field, catalyzing innovation and establishing new performance benchmarks. Image registration, a foundati…

eess.IV2026

Large-Scale Deployment and Analytical Implications of Structured Quality Control in Diffusion Magnetic Resonance Imaging

Michael E. Kim, Chenyu Gao, Karthik Ramadass +17

Purpose: Diffusion MRI (dMRI) provides a diverse set of quantitative measures and derived datatypes to assess white matter microstructure and macrostructure. Coupled with the incre…

cs.AI2026

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models

Siqi Luo, Jianghan Shen, Yi Xin +9

Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One prima…

cs.CV2026

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

Jiayang Li, Shuo Cao, Xiaohui Li +6

In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. H…