A Survey of AI Text-to-Image and AI Text-to-Video Generators
arXiv:2311.06329 · doi:10.1109/AIRC57904.2023.10303174
Abstract
Text-to-Image and Text-to-Video AI generation models are revolutionary technologies that use deep learning and natural language processing (NLP) techniques to create images and videos from textual descriptions. This paper investigates cutting-edge approaches in the discipline of Text-to-Image and Text-to-Video AI generations. The survey provides an overview of the existing literature as well as an analysis of the approaches used in various studies. It covers data preprocessing techniques, neural network types, and evaluation metrics used in the field. In addition, the paper discusses the challenges and limitations of Text-to-Image and Text-to-Video AI generations, as well as future research directions. Overall, these models have promising potential for a wide range of applications such as video production, content creation, and digital marketing.
4 pages, 2 tables, 4th International Conference on Artificial Intelligence, Robotics and Control (AIRC 2023)
References in corpus (9)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
- Imagen Video: High Definition Video Generation with Diffusion Models
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- Multimodal datasets: misogyny, pornography, and malignant stereotypes
- CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
- CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
- Phenaki: Variable Length Video Generation From Open Domain Textual Description
- Text-to-Image Generation with Attention Based Recurrent Neural Networks