2 citations · 6 across the 12 of their papers we have counts for
11 papers · 1 filter
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
Ashkan Taghipour, Morteza Ghahremani, Zinuo Li +3
Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally…
AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
Xian Zhang, Zexi Wu, Zinuo Li +5
Understanding long-form videos remains a significant challenge for vision--language models (VLMs) due to their extensive temporal length and high information density. Most current…
LatentMove: Towards Complex Human Movement Video Generation
Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun +5
Image-to-video (I2V) generation seeks to produce realistic motion sequences from a single reference image. Although recent methods exhibit strong temporal consistency, they often s…
Layer-Wise Feature Metric of Semantic-Pixel Matching for Few-Shot Learning
Hao Tang, Junhao Lu, Guoheng Huang +5
In Few-Shot Learning (FSL), traditional metric-based approaches often rely on global metrics to compute similarity. However, in natural scenes, the spatial arrangement of key insta…
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
Ashkan Taghipour, Morteza Ghahremani, Mohammed Bennamoun +4
This paper investigates the role of CLIP image embeddings within the Stable Video Diffusion (SVD) framework, focusing on their impact on video generation quality and computational…
CLIP Guided Image-perceptive Prompt Learning for Image Enhancement
Weiwen Chen, Qiuhong Ke, Zinuo Li
Image enhancement is a significant research area in the fields of computer vision and image processing. In recent years, many learning-based methods for image enhancement have been…