3 citations · 4 across the 3 of their papers we have counts for
4 papers
Rethinking Video Tokenization: A Conditioned Diffusion-based Approach
Nianzu Yang, Pandeng Li, Liming Zhao +8
Existing video tokenizers typically use the traditional Variational Autoencoder (VAE) architecture for video compression and reconstruction. However, to achieve good performance, i…
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
Zhihang Liu, Chen-Wei Xie, Bin Wen +9
Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics…
Motion Control for Enhanced Complex Action Video Generation
Qiang Zhou, Shaofeng Zhang, Nianzu Yang +2
Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to p…
Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation
Chao Chen, Haoyu Geng, Nianzu Yang +4
User interests are usually dynamic in the real world, which poses both theoretical and practical challenges for learning accurate preferences from rich behavior data. Among existin…