activity
20182026
most citedFrom Screens to Scenes: A Survey of Embodied AI in Healthcare

36 citations · 143 across the 63 of their papers we have counts for

collaborators
Showing cs.CVShow all

49 papers · 1 filter

cs.CV2026

Distilling Physical Priors into Streaming World Models

Liangliang Zhao, Junying Wang, Danni Yang +5

Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical c…

cs.CV2026

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

Jiayang Li, Shuo Cao, Xiaohui Li +6

In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. H…

cs.CV2026

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

Changyao Tian, Danni Yang, Guanzhou Chen +26

Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade-offs between maintaining strong semantic comprehension and acq…

cs.CV2026

Accelerating Masked Image Generation by Learning Controlled Latent Dynamics

Kaiwen Zhu, Quansheng Zeng, Yuandong Pu +9

Masked Image Generation Models (MIGMs) have achieved great success, yet their efficiency is hampered by the multiple steps of bi-directional attention. In fact, there exists notabl…

cs.CV2026

HSD: Training-Free Acceleration for Document Parsing Vision-Language Models with Hierarchical Speculative Decoding

Wenhui Liao, Hongliang Li, Pengyu Xie +15

Document parsing is a fundamental task in multimodal understanding, supporting a wide range of downstream applications such as information extraction and intelligent document analy…

cs.CV2026

Toward Generalizable Deblurring: Leveraging Massive Blur Priors with Linear Attention for Real-World Scenarios

Yuanting Gao, Shuo Cao, Xiaohui Li +3

Image deblurring has advanced rapidly with deep learning, yet most methods exhibit poor generalization beyond their training datasets, with performance dropping significantly in re…