activity
20192025
most citedVideoComposer: Compositional Video Synthesis with Motion Controllability

43 citations · 159 across the 25 of their papers we have counts for

collaborators
Showing cs.CVShow all

26 papers · 1 filter

cs.CV2025

HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation

Xiang Wang, Zhifei Zhang, He Zhang +11

Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanc…

cs.CV20232 cited

Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer Learning

Zhiwu Qing, Shiwei Zhang, Ziyuan Huang +4

Recently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding spatial contents, but naively transferring such models to…

cs.CV2023

Towards Real-World Visual Tracking with Temporal Contexts

Ziang Cao, Ziyuan Huang, Liang Pan +3

Visual tracking has made significant improvements in the past few decades. Most existing state-of-the-art trackers 1) merely aim for performance in ideal conditions while overlooki…

cs.CV20237 cited

Temporally-Adaptive Models for Efficient Video Understanding

Ziyuan Huang, Shiwei Zhang, Liang Pan +4

Spatial convolutions are extensively used in numerous deep video models. It fundamentally assumes spatio-temporal invariance, i.e., using shared weights for every location in diffe…

cs.CV202343 cited

VideoComposer: Compositional Video Synthesis with Motion Controllability

Xiang Wang, Hangjie Yuan, Shiwei Zhang +6

The pursuit of controllability as a higher standard of visual content creation has yielded remarkable progress in customizable image synthesis. However, achieving controllable vide…

cs.CV20233 cited

MoLo: Motion-augmented Long-short Contrastive Learning for Few-shot Action Recognition

Xiang Wang, Shiwei Zhang, Zhiwu Qing +4

Current state-of-the-art approaches for few-shot action recognition achieve promising performance by conducting frame-level matching on learned visual features. However, they gener…