most citedAdaptive Frequency Filters As Efficient Global Token Mixers

10 citations · 12 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2024

Text Grouping Adapter: Adapting Pre-trained Text Detector for Layout Analysis

Tianci Bi, Xiaoyi Zhang, Zhizheng Zhang +4

Significant progress has been made in scene text detection models since the rise of deep learning, but scene text layout analysis, which aims to group detected text instances as pa…

cs.CV2024

RelationVLM: Making Large Vision-Language Models Understand Visual Relations

Zhipeng Huang, Zhizheng Zhang, Zheng-Jun Zha +2

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very…

cs.CV2024

Slot-VLM: SlowFast Slots for Video-Language Modeling

Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2

Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development…

cs.CV2023

Long Video Understanding with Learnable Retrieval in Video-Language Models

Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2

The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understandi…

cs.CV20232 cited

Reinforced UI Instruction Grounding: Towards a Generic UI Task Automation API

Zhizheng Zhang, Wenxuan Xie, Xiaoyi Zhang +1

Recent popularity of Large Language Models (LLMs) has opened countless possibilities in automating numerous AI tasks by connecting LLMs to various domain-specific models or APIs, w…

cs.CV202310 cited

Adaptive Frequency Filters As Efficient Global Token Mixers

Zhipeng Huang, Zhizheng Zhang, Cuiling Lan +3

Recent vision transformers, large-kernel CNNs and MLPs have attained remarkable successes in broad vision tasks thanks to their effective information fusion in the global scope. Ho…