activity
20182026
most citedPyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

42 citations · 46 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2025

Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation

Inclusion AI, :, Bowen Ma +73

We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which on…

cs.CV2024

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts

Honglin Li, Yuting Gao, Chenglu Zhu +3

Multimodal large language models (MLLMs) are closing the gap to human visual perception capability rapidly, while, still lag behind on attending to subtle images details or locatin…

cs.CV202242 cited

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining

Yuting Gao, Jinfeng Liu, Zihan Xu +4

Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from t…

cs.CV2021

An Empirical Study and Analysis on Open-Set Semi-Supervised Learning

Huixiang Luo, Hao Cheng, Fanxu Meng +4

Pseudo-labeling (PL) and Data Augmentation-based Consistency Training (DACT) are two approaches widely used in Semi-Supervised Learning (SSL) methods. These methods exhibit great p…

cs.CV2020

Association: Remind Your GAN not to Forget

Yi Gu, Jie Li, Yuting Gao +5

Neural networks are susceptible to catastrophic forgetting. They fail to preserve previously acquired knowledge when adapting to new tasks. Inspired by human associative memory sys…

cs.CV2020

Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning

Jinpeng Wang, Yuting Gao, Ke Li +7

Self-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some…