papers

Publications (6)

cs.CV2025

SkyReels-V2: Infinite-length Film Generative Model

Guibin Chen, Dixuan Lin, Jiangping Yang +22

Recent advances in video generation have been driven by diffusion models and autoregressive frameworks, yet critical challenges persist in harmonizing prompt adherence, visual qual…

cs.CV2026

OmniHands: Towards Robust 4D Hand Mesh Recovery via A Versatile Transformer

Dixuan Lin, Yuxiang Zhang, Mengcheng Li +5

In this paper, we introduce OmniHands, a universal approach to recovering interactive hand meshes and their relative movement from monocular or multi-view inputs. Our approach addr…

cs.CV2026

SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model

Guibin Chen, Dixuan Lin, Jiangping Yang +48

SkyReels V4 is a unified multi modal video foundation model for joint video audio generation, inpainting, and editing. The model adopts a dual stream Multimodal Diffusion Transform…

cs.CV2026

SkyReels-V3 Technique Report

Debang Li, Zhengcong Fei, Tuanhui Li +19

Video generation serves as a cornerstone for building world models, where multimodal contextual inference stands as the defining test of capability. In this end, we present SkyReel…

cs.CV2025

Zero-shot Reconstruction of In-Scene Object Manipulation from Video

Dixuan Lin, Tianyou Wang, Zhuoyang Pan +3

We build the first system to address the problem of reconstructing in-scene object manipulation from a monocular RGB video. It is challenging due to ill-posed scene reconstruction,…

cs.CV2023

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Dixuan Lin, Yixing Peng, Jingke Meng +1

Text-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed…