activity
20202023
most citedTinyViT: Fast Pretraining Distillation for Small Vision Transformers

22 citations · 92 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CV20231 cited

MovieFactory: Automatic Movie Creation from Text using Large Generative Models for Language and Images

Junchen Zhu, Huan Yang, Huiguo He +6

In this paper, we present MovieFactory, a powerful framework to generate cinematic-picture (30721280), film-style (multi-scene), and multi-modality (sounding) movies on the…

cs.RO20239 cited

AlphaBlock: Embodied Finetuning for Vision-Language Reasoning in Robot Manipulation

Chuhao Jin, Wenhui Tan, Jiange Yang +4

We propose a novel framework for learning high-level cognitive capabilities in robot manipulation tasks, such as making a smiley face using building blocks. These tasks often invol…

cs.CV20231 cited

NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation

Shengming Yin, Chenfei Wu, Huan Yang +13

In this paper, we propose NUWA-XL, a novel Diffusion over Diffusion architecture for eXtremely Long video generation. Most current work generates long videos segment by segment seq…

cs.CV202317 cited

Unified Multi-Modal Latent Diffusion for Joint Subject and Text Conditional Image Generation

Yiyang Ma, Huan Yang, Wenjing Wang +2

Language-guided image generation has achieved great success nowadays by using diffusion models. However, texts can be less detailed to describe highly-specific subjects such as a p…

cs.CV20222 cited

Exploring Anchor-based Detection for Ego4D Natural Language Query

Sipeng Zheng, Qi Zhang, Bei Liu +2

In this paper we provide the technique report of Ego4D natural language query challenge in CVPR 2022. Natural language query task is challenging due to the requirement of comprehen…

cs.CV2022

GRIT-VLP: Grouped Mini-batch Sampling for Efficient Vision and Language Pre-training

Jaeseok Byun, Taebaek Hwang, Jianlong Fu +1

Most of the currently existing vision and language pre-training (VLP) methods have mainly focused on how to extract and align vision and text features. In contrast to the mainstrea…