activity
20142025
most citedDeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection

134 citations · 565 across the 52 of their papers we have counts for

collaborators
Showing cs.CVShow all

36 papers · 1 filter

cs.CV20246 cited

MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Kaining Ying, Fanqing Meng, Jin Wang +19

Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimod…

cs.CV2024

PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Junsong Chen, Chongjian Ge, Enze Xie +7

In this paper, we introduce PixArt-Σ, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution. PixArt-Σrepresents a significant advancement over…

cs.CV202417 cited

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Zhe Chen, Jiannan Wu, Wenhai Wang +12

The exponential growth of large language models (LLMs) has opened up numerous possibilities for multimodal AGI systems. However, the progress in vision and vision-language foundati…

cs.CV202310 cited

Flow-Based Feature Fusion for Vehicle-Infrastructure Cooperative 3D Object Detection

Haibao Yu, Yingjuan Tang, Enze Xie +3

Cooperatively utilizing both ego-vehicle and infrastructure sensor data can significantly enhance autonomous driving perception abilities. However, the uncertain temporal asynchron…

cs.CV2023

RIGID: Recurrent GAN Inversion and Editing of Real Face Videos

Yangyang Xu, Shengfeng He, Kwan-Yee K. Wong +1

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired in…

cs.CV2023

Exploring Transformers for Open-world Instance Segmentation

Jiannan Wu, Yi Jiang, Bin Yan +3

Open-world instance segmentation is a rising task, which aims to segment all objects in the image by learning from a limited number of base-category objects. This task is challengi…