activity
20242026
most citedM2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

4 citations · 4 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

KlingAvatar 2.0 Technical Report

Kling Team, Jialu Chen, Yikang Ding +25

Avatar video generation models have achieved remarkable progress in recent years. However, prior work exhibits limited efficiency in generating long-duration high-resolution videos…

cs.CV2025

Decoupling Complexity from Scale in Latent Diffusion Model

Tianxiong Zhong, Xingye Tian, Xuebo Wang +3

Existing latent diffusion models typically couple scale with content complexity, using more latent tokens to represent higher-resolution images or higher-frame rate videos. However…

cs.CV2025

VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption

Tianxiong Zhong, Xingye Tian, Boyuan Jiang +4

Modern video generation frameworks based on Latent Diffusion Models suffer from inefficiencies in tokenization due to the Frame-Proportional Information Assumption. Existing tokeni…

cs.CV2024

VIVID-10M: A Dataset and Baseline for Versatile and Interactive Video Local Editing

Jiahao Hu, Tianxiong Zhong, Xuebo Wang +5

Diffusion-based image editing models have made remarkable progress in recent years. However, achieving high-quality video editing remains a significant challenge. One major hurdle…

cs.CV2024

Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

Qiuheng Wang, Yukai Shi, Jiarong Ou +10

With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the perfo…