activity
20242026
collaborators

7 papers

cs.CV2026

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Guozhen Zhang, Xuerui Qiu, Yutao Cui +11

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDR…

cs.CV2026

Arbitrary Generative Video Interpolation

Guozhen Zhang, Haiguang Wang, Chunyu Wang +3

Video frame interpolation (VFI), which generates intermediate frames from given start and end frames, has become a fundamental function in video generation applications. However, e…

cs.CV2025

Motion-Aware Generative Frame Interpolation

Guozhen Zhang, Yuhan Zhu, Yutao Cui +3

Flow-based frame interpolation methods ensure motion stability through estimated intermediate flow but often introduce severe artifacts in complex motion regions. Recent generative…

cs.CV2024

VFIMamba: Video Frame Interpolation with State Space Models

Guozhen Zhang, Chunxu Liu, Yutao Cui +3

Inter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based model…

cs.CV2024

Sparse Global Matching for Video Frame Interpolation with Large Motion

Chunxu Liu, Guozhen Zhang, Rui Zhao +1

Large motion poses a critical challenge in Video Frame Interpolation (VFI) task. Existing methods are often constrained by limited receptive fields, resulting in sub-optimal perfor…

cs.CV2024

Dynamic and Compressive Adaptation of Transformers From Images to Videos

Guozhen Zhang, Jingyu Liu, Shengming Cao +4

Recently, the remarkable success of pre-trained Vision Transformers (ViTs) from image-text matching has sparked an interest in image-to-video adaptation. However, most current appr…