From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
arXiv:2312.10868 · doi:10.3390/technologies13020051
Abstract
This comprehensive survey explored the evolving landscape of generative Artificial Intelligence (AI), with a specific focus on the transformative impacts of Mixture of Experts (MoE), multimodal learning, and the speculated advancements towards Artificial General Intelligence (AGI). It critically examined the current state and future trajectory of generative Artificial Intelligence (AI), exploring how innovations like Google's Gemini and the anticipated OpenAI Q* project are reshaping research priorities and applications across various domains, including an impact analysis on the generative AI research taxonomy. It assessed the computational challenges, scalability, and real-world implications of these technologies while highlighting their potential in driving significant progress in fields like healthcare, finance, and education. It also addressed the emerging academic challenges posed by the proliferation of both AI-themed and AI-generated preprints, examining their impact on the peer-review process and scholarly communication. The study highlighted the importance of incorporating ethical and human-centric methods in AI development, ensuring alignment with societal norms and welfare, and outlined a strategy for future AI research that focuses on a balanced and conscientious use of MoE, multimodality, and AGI in generative AI.
30 pages
References in corpus (24)
- Survey of Hallucination in Natural Language Generation
- A novel time-frequency Transformer based on self-attention mechanism and its application in fault diagnosis of rolling bearings
- Exploration in Deep Reinforcement Learning: A Survey
- TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech
- Designing Creative AI Partners with COFI: A Framework for Modeling Interaction in Human-AI Co-Creative Systems
- Towards artificial general intelligence via a multimodal foundation model
- Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions
- Large-scale Text-to-Image Generation Models for Visual Artists' Creative Works
- GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
- Self-Supervised Learning for Videos: A Survey
- A Prescriptive Learning Analytics Framework: Beyond Predictive Modelling and onto Explainable AI with Prescriptive Analytics and ChatGPT
- On Recurrent Neural Networks for learning-based control: recent results and ideas for future developments
- From COBIT to ISO 42001: Evaluating Cybersecurity Frameworks for Opportunities, Risks, and Regulatory Compliance in Commercializing Large Language Models
- Can Large Language Models Reason and Plan?
- GPT-NeoX-20B: An Open-Source Autoregressive Language Model
- Machine Learning in NextG Networks via Generative Adversarial Networks
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
- Privately Fine-Tuning Large Language Models with Differential Privacy
- Human-Centric Multimodal Machine Learning: Recent Advances and Testbed on AI-based Recruitment
- Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias
- Three lines of defense against risks from AI
- MViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design
- TALM: Tool Augmented Language Models
- Improving Policy Optimization with Generalist-Specialist Learning