activity
20152024
most citedMMDetection: Open MMLab Detection Toolbox and Benchmark

794 citations · 1.3k across the 25 of their papers we have counts for

collaborators
Showing cs.CVShow all

43 papers · 1 filter

cs.CV2024

Learning 1D Causal Visual Representation with De-focus Attention Networks

Chenxin Tao, Xizhou Zhu, Shiqian Su +8

Modality differences have led to the development of heterogeneous architectures for vision and language models. While images typically require 2D non-causal modeling, texts utilize…

cs.CV2024

Parameter-Inverted Image Pyramid Networks

Xizhou Zhu, Xue Yang, Zhaokai Wang +6

Image pyramids are commonly used in modern computer vision tasks to obtain multi-scale features for precise understanding of images. However, image pyramids process multiple resolu…

cs.CV2024

Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

Yuchen Duan, Weiyun Wang, Zhe Chen +7

Transformers have revolutionized computer vision and natural language processing, but their high computational complexity limits their application in high-resolution image processi…

cs.CV2024

The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

Weiyun Wang, Yiming Ren, Haowen Luo +9

We present the All-Seeing Project V2: a new model and dataset designed for understanding object relations in images. Specifically, we propose the All-Seeing Model V2 (ASMv2) that i…

cs.CV202417 cited

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Zhe Chen, Jiannan Wu, Wenhai Wang +12

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundati…

cs.CV20249 cited

Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications

Yuwen Xiong, Zhiqi Li, Yuntao Chen +10

We introduce Deformable Convolution v4 (DCNv4), a highly efficient and effective operator designed for a broad spectrum of vision applications. DCNv4 addresses the limitations of i…