GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text
arXiv:2308.06911 · doi:10.1016/j.compbiomed.2024.108073
Abstract
Large language models have made significant strides in natural language processing, enabling innovative applications in molecular science by processing textual representations of molecules. However, most existing language models cannot capture the rich information with complex molecular structures or images. In this paper, we introduce GIT-Mol, a multi-modal large language model that integrates the Graph, Image, and Text information. To facilitate the integration of multi-modal molecular data, we propose GIT-Former, a novel architecture that is capable of aligning all modalities into a unified latent space. We achieve a 5%-10% accuracy increase in properties prediction and a 20.2% boost in molecule generation validity compared to the baselines. With the any-to-language molecular translation strategy, our model has the potential to perform more downstream tasks, such as compound name recognition and chemical reaction prediction.
The article has been accepted by Computers in Biology and Medicine, with 14 pages and 4 figures
References in corpus (8)
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
- Language Is Not All You Need: Aligning Perception with Language Models
- Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
- Empowering Molecule Discovery for Molecule-Caption Translation with Large Language Models: A ChatGPT Perspective
- A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language
- One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale
- Foundation Transformers