10 citations · 12 across the 4 of their papers we have counts for
6 papers
Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis
Tianci Bi, Xiaoyi Zhang, Zhizheng Zhang +4
Significant progress has been made in scene text detection models since the rise of deep learning, but scene text layout analysis, which aims to group detected text instances as pa…
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
Zhipeng Huang, Zhizheng Zhang, Zheng-Jun Zha +2
The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very…
Slot-VLM: SlowFast Slots for Video-Language Modeling
Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2
Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development…
Long Video Understanding with Learnable Retrieval in Video-Language Models
Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2
The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understandi…
Reinforced UI Instruction Grounding: Towards a Generic UI Task Automation API
Zhizheng Zhang, Wenxuan Xie, Xiaoyi Zhang +1
Recent popularity of Large Language Models (LLMs) has opened countless possibilities in automating numerous AI tasks by connecting LLMs to various domain-specific models or APIs, w…
Adaptive Frequency Filters As Efficient Global Token Mixers
Zhipeng Huang, Zhizheng Zhang, Cuiling Lan +3
Recent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. Ho…