42 citations · 46 across the 8 of their papers we have counts for
8 papers · 1 filter
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
Inclusion AI, :, Bowen Ma +73
We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which on…
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
Honglin Li, Yuting Gao, Chenglu Zhu +3
Multimodal large language models (MLLMs) are closing the gap to human visual perception capability rapidly, while, still lag behind on attending to subtle images details or locatin…
PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining
Yuting Gao, Jinfeng Liu, Zihan Xu +4
Large-scale vision-language pre-training has achieved promising results on downstream tasks. Existing methods highly rely on the assumption that the image-text pairs crawled from t…
An Empirical Study and Analysis on Open-Set Semi-Supervised Learning
Huixiang Luo, Hao Cheng, Fanxu Meng +4
Pseudo-labeling (PL) and Data Augmentation-based Consistency Training (DACT) are two approaches widely used in Semi-Supervised Learning (SSL) methods. These methods exhibit great p…
Association: Remind Your GAN not to Forget
Yi Gu, Jie Li, Yuting Gao +5
Neural networks are susceptible to catastrophic forgetting. They fail to preserve previously acquired knowledge when adapting to new tasks. Inspired by human associative memory sys…
Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
Jinpeng Wang, Yuting Gao, Ke Li +7
Self-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some…