10 papers · 1 filter
Bridging Video Understanding and Generation in a Unified Framework
Yuqi Wang, Runyi Li, Ruoyu Feng +3
Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexp…
Video-Mirai: Autoregressive Video Diffusion Models Need Foresight
Yonghao Yu, Lang Huang, Runyi Li +2
Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment tha…
Mirai: Autoregressive Visual Generation Needs Foresight
Yonghao Yu, Lang Huang, Zerun Wang +2
Autoregressive (AR) visual generators model images as sequences of discrete tokens and are trained with a next-token likelihood objective. This strict causal supervision optimizes…
GaussianSeal: Rooting Adaptive Watermarks for 3D Gaussian Generation Model
Runyi Li, Xuanyu Zhang, Chuhan Tong +2
With the advancement of AIGC technologies, the modalities generated by models have expanded from images and videos to 3D objects, leading to an increasing number of works focused o…
LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
Runyi Li, Bin Chen, Jian Zhang +1
Blind face restoration from low-quality images is a challenging task that requires not only high-fidelity image reconstruction, but also preservation of facial identity. Although d…
FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models
Zhipei Xu, Xuanyu Zhang, Runyi Li +3
The rapid development of generative AI is a double-edged sword, which not only facilitates content creation but also makes image manipulation easier and more difficult to detect. A…