8 citations · 23 across the 45 of their papers we have counts for
8 papers · 1 filter
Make-A-Storyboard: A General Framework for Storyboard with Disentangled and Merged Control
Sitong Su, Litao Guo, Lianli Gao +2
Story Visualization aims to generate images aligned with story prompts, reflecting the coherence of storybooks through visual consistency among characters and scenes.Whereas curren…
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
Sitong Su, Jianzhi Liu, Lianli Gao +1
Recently Text-to-Video (T2V) synthesis has undergone a breakthrough by training transformers or diffusion models on large-scale datasets. Nevertheless, inferring such large models…
ALF: Adaptive Label Finetuning for Scene Graph Generation
Qishen Chen, Jianzhi Liu, Xinyu Lyu +3
Scene Graph Generation (SGG) endeavors to predict the relationships between subjects and objects in a given image. Nevertheless, the long-tail distribution of relations often leads…
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
Kaipeng Fang, Jingkuan Song, Lianli Gao +4
The goal of Universal Cross-Domain Retrieval (UCDR) is to achieve robust performance in generalized test scenarios, wherein data may belong to strictly unknown domains and categori…
MotionZero:Exploiting Motion Priors for Zero-shot Text-to-Video Generation
Sitong Su, Litao Guo, Lianli Gao +2
Zero-shot Text-to-Video synthesis generates videos based on prompts without any videos. Without motion information from videos, motion priors implied in prompts are vital guidance.…
BatchNorm-based Weakly Supervised Video Anomaly Detection
Yixuan Zhou, Yi Qu, Xing Xu +3
In weakly supervised video anomaly detection (WVAD), where only video-level labels indicating the presence or absence of abnormal events are available, the primary challenge arises…