8 papers
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
Leyi Wu, Yifan Zhao, Jinjie Zhang +11
Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception under real conditions is essent…
CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning
Tianshuo Xu, Tiantian Hong, Zhifei Chen +2
Page-level calligraphy synthesis requires balancing glyph precision with layout composition. Existing character models lack spatial context, while page-level methods often compromi…
Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics
Tianshuo Xu, Zhifei Chen, Leyi Wu +2
The ultimate goal of video generation is to satisfy a fundamental trilemma: achieving high visual quality, maintaining rigorous physical consistency, and enabling precise controlla…
UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
Tianshuo Xu, Kai Wang, Zhifei Chen +4
Computational replication of Chinese calligraphy remains challenging. Existing methods falter, either creating high-quality isolated characters while ignoring page-level aesthetics…
A Mechanistic View on Video Generation as World Models: State and Dynamics
Luozhou Wang, Zhifei Chen, Yihua Du +11
Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateles…
FlexPainter: Flexible and Multi-View Consistent Texture Generation
Dongyu Yan, Leyi Wu, Jiantao Lin +7
Texture map production is an important part of 3D modeling and determines the rendering quality. Recently, diffusion-based methods have opened a new way for texture generation. How…