activity
20152022
most citedA Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language

42 citations · 144 across the 13 of their papers we have counts for

collaborators

25 papers

cs.LG202242 cited

A Molecular Multimodal Foundation Model Associating Molecule Graphs with Natural Language

Bing Su, Dazhao Du, Zhao Yang +6

Although artificial intelligence (AI) has made significant progress in understanding molecules in a wide range of fields, existing models generally acquire the single cognitive abi…

cs.CV20222 cited

COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval

Haoyu Lu, Nanyi Fei, Yuqi Huo +3

Large-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recentl…

cs.CV20213 cited

HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers

Mingyu Ding, Xiaochen Lian, Linjie Yang +4

High-resolution representations (HR) are essential for dense prediction tasks such as segmentation, detection, and pose estimation. Learning HR representations is typically ignored…

cs.AI202115 cited

Pre-Trained Models: Past, Present and Future

Xu Han, Zhengyan Zhang, Ning Ding +21

Large-scale pre-trained models (PTMs) such as BERT and GPT have recently achieved great success and become a milestone in the field of artificial intelligence (AI). Owing to sophis…

cs.CV2021

WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training

Yuqi Huo, Manli Zhang, Guangzhen Liu +32

Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…

cs.CV20218 cited

Contrastive Prototype Learning with Augmented Embeddings for Few-Shot Learning

Yizhao Gao, Nanyi Fei, Guangzhen Liu +3

Most recent few-shot learning (FSL) methods are based on meta-learning with episodic training. In each meta-training episode, a discriminative feature embedding and/or classifier a…