9 papers
Bridging Video Understanding and Generation in a Unified Framework
Yuqi Wang, Runyi Li, Ruoyu Feng +3
Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexp…
Video-Mirai: Autoregressive Video Diffusion Models Need Foresight
Yonghao Yu, Lang Huang, Runyi Li +2
Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segment becomes a commitment tha…
Mirai: Autoregressive Visual Generation Needs Foresight
Yonghao Yu, Lang Huang, Zerun Wang +2
Autoregressive (AR) visual generators model images as sequences of discrete tokens and are trained with a next-token likelihood objective. This strict causal supervision optimizes…
RealOSR: Latent Guidance Boosts Diffusion-based Real-world Omnidirectional Image Super-Resolutions
Xuhan Sheng, Runyi Li, Bin Chen +3
Omnidirectional image super-resolution (ODISR) aims to upscale low-resolution (LR) omnidirectional images (ODIs) to high-resolution (HR), catering to the growing demand for detaile…
GaussianSeal: Rooting Adaptive Watermarks for 3D Gaussian Generation Model
Runyi Li, Xuanyu Zhang, Chuhan Tong +2
With the advancement of AIGC technologies, the modalities generated by models have expanded from images and videos to 3D objects, leading to an increasing number of works focused o…
LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
Runyi Li, Bin Chen, Jian Zhang +1
Blind face restoration from low-quality (LQ) images is a challenging task that requires not only high-fidelity image reconstruction but also the preservation of facial identity. Wh…