activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Guozhen Zhang, Xuerui Qiu, Yutao Cui +11

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDR…

cs.CV2026

UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal Interactions

Guozhen Zhang, Zixiang Zhou, Teng Hu +6

Due to the lack of effective cross-modal modeling, existing open-source audio-video generation methods often exhibit compromised lip synchronization and insufficient semantic consi…

cs.CV2026

Arbitrary Generative Video Interpolation

Guozhen Zhang, Haiguang Wang, Chunyu Wang +3

Video frame interpolation (VFI), which generates intermediate frames from given start and end frames, has become a fundamental function in video generation applications. However, e…

cs.CV2025

Motion-Aware Generative Frame Interpolation

Guozhen Zhang, Yuhan Zhu, Yutao Cui +3

Flow-based frame interpolation methods ensure motion stability through estimated intermediate flow but often introduce severe artifacts in complex motion regions. Recent generative…

cs.CV2024

VFIMamba: Video Frame Interpolation with State Space Models

Guozhen Zhang, Chunxu Liu, Yutao Cui +3

Inter-frame modeling is pivotal in generating intermediate frames for video frame interpolation (VFI). Current approaches predominantly rely on convolution or attention-based model…

cs.CV2024

Sparse Global Matching for Video Frame Interpolation with Large Motion

Chunxu Liu, Guozhen Zhang, Rui Zhao +1

Large motion poses a critical challenge in Video Frame Interpolation (VFI) task. Existing methods are often constrained by limited receptive fields, resulting in sub-optimal perfor…