2 papers
cs.CV2024
Blended Latent Diffusion under Attention Control for Real-World Video Editing
Deyin Liu, Lin Yuanbo Wu, Xianghua Xie
Due to lack of fully publicly available text-to-video models, current video editing methods tend to build on pre-trained text-to-image generation models, however, they still face g…
cs.CV2024
In-context Prompt Learning for Test-time Vision Recognition with Frozen Vision-language Model
Junhui Yin, Xinyu Zhang, Lin Wu +1
Current pre-trained vision-language models, such as CLIP, have demonstrated remarkable zero-shot generalization capabilities across various downstream tasks. However, their perform…