6 citations · 8 across the 4 of their papers we have counts for
6 papers
Dynamic Token Reduction during Generation for Vision Language Models
Xiaoyu Liang, Chaofeng Guan, Jiaying Lu +3
Vision-Language Models (VLMs) have achieved notable success in multimodal tasks but face practical limitations due to the quadratic complexity of decoder attention mechanisms and a…
ST: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token Trimming
Jiedong Zhuang, Lu Lu, Ming Dai +4
Multimodal large language models (MLLMs) enhance their perceptual capabilities by integrating visual and textual information. However, processing the massive number of visual token…
FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance
Jiedong Zhuang, Jiaqi Hu, Lianrui Mu +4
CLIP has achieved impressive zero-shot performance after pre-training on a large-scale dataset consisting of paired image-text data. Previous works have utilized CLIP by incorporat…
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
Jianhong Bai, Tianyu He, Yuchi Wang +4
Recent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.…
On the Effectiveness of Out-of-Distribution Data in Self-Supervised Long-Tail Learning
Jianhong Bai, Zuozhu Liu, Hualiang Wang +4
Though Self-supervised learning (SSL) has been widely studied as a promising technique for representation learning, it doesn't generalize well on long-tailed datasets due to the ma…
Towards Calibrated Hyper-Sphere Representation via Distribution Overlap Coefficient for Long-tailed Learning
Hualiang Wang, Siming Fu, Xiaoxuan He +3
Long-tailed learning aims to tackle the crucial challenge that head classes dominate the training procedure under severe class imbalance in real-world scenarios. However, little at…