most citedImproved Video VAE for Latent Video Diffusion Model

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.LG2025

Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment

Yawen Shao, Jie Xiao, Kai Zhu +4

Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO…

cs.CV2025

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

Jian Yang, Dacheng Yin, Xiaoxuan He +6

Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. W…

cs.CV2025

VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization

Zixun Fang, Zhiheng Liu, Kai Zhu +5

Video colorization aims to transform grayscale videos into vivid color representations while maintaining temporal consistency and structural integrity. Existing video colorization…

cs.CV2024

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

Fan Lu, Wei Wu, Kecheng Zheng +7

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have dev…

cs.CV20241 cited

Improved Video VAE for Latent Video Diffusion Model

Pingyu Wu, Kai Zhu, Yu Liu +4

Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation…