1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
Vision as LoRA
Han Wang, Yongjie Ye, Bingru Li +5
We introduce Vision as LoRA (VoRA), a novel paradigm for transforming an LLM into an MLLM. Unlike prevalent MLLM architectures that rely on external vision modules for vision encod…
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
Han Wang, Yuxiang Nie, Yongjie Ye +6
The application of Large Vision-Language Models (LVLMs) for analyzing images and videos is an exciting and rapidly evolving field. In recent years, we've seen significant growth in…
Elysium: Exploring Object-level Perception in Videos via MLLM
Han Wang, Yanjie Wang, Yongjie Ye +2
Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking…
GloTSFormer: Global Video Text Spotting Transformer
Han Wang, Yanjie Wang, Yang Li +1
Video Text Spotting (VTS) is a fundamental visual task that aims to predict the trajectories and content of texts in a video. Previous works usually conduct local associations and…