3 citations · 5 across the 8 of their papers we have counts for
8 papers · 1 filter
LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
Lianyu Hu, Fanhua Shang, Wei Feng +1
In this paper, we introduce LightVLM, a simple but effective method that can be seamlessly deployed upon existing Vision-Language Models (VLMs) to greatly accelerate the inference…
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
Lianyu Hu, Liqing Gao, Fanhua Shang +2
Recent methods have made notable progress in accelerating Large Vision-Language Models (LVLMs) by exploiting the inherent redundancy in visual inputs. Most existing approaches, how…
Deep Correlated Prompting for Visual Recognition with Missing Modalities
Lianyu Hu, Tongkai Shi, Wei Feng +2
Large-scale multimodal models have shown excellent performance over a series of tasks powered by the large corpus of paired multimodal training data. Generally, they are always ass…
Pose-Guided Fine-Grained Sign Language Video Generation
Tongkai Shi, Lianyu Hu, Fanhua Shang +3
Sign language videos are an important medium for spreading and learning sign language. However, most existing human image synthesis methods produce sign language images with detail…
CorrNet+: Sign Language Recognition and Translation via Spatial-Temporal Correlation
Lianyu Hu, Wei Feng, Liqing Gao +2
In sign language, the conveyance of human body trajectories predominantly relies upon the coordinated movements of hands and facial expressions across successive frames. Despite th…
Improving Continuous Sign Language Recognition with Adapted Image Models
Lianyu Hu, Tongkai Shi, Liqing Gao +2
The increase of web-scale weakly labelled image-text pairs have greatly facilitated the development of large-scale vision-language models (e.g., CLIP), which have shown impressive…