activity
20212023
most citedVisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

131 citations · 233 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CV20231 cited

FlowFormer: A Transformer Architecture and Its Masked Cost Volume Autoencoding for Optical Flow

Zhaoyang Huang, Xiaoyu Shi, Chao Zhang +6

This paper introduces a novel transformer-based network architecture, FlowFormer, along with the Masked Cost Volume AutoEncoding (MCVA) for pretraining it to tackle the problem of…

cs.CV202314 cited

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Zhaoyang Liu, Yinan He, Wenhai Wang +17

We present an interactive visual framework named InternGPT, or iGPT for short. The framework integrates chatbots that have planning and reasoning capabilities, such as ChatGPT, wit…

cs.AI202324 cited

Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory

Xizhou Zhu, Yuntao Chen, Hao Tian +10

The captivating realm of Minecraft has attracted substantial research interest in recent years, serving as a rich platform for developing intelligent agents capable of functioning…

cs.CV2023131 cited

VisionLLM: Large Language Model is also an Open-Ended Decoder for Vision-Centric Tasks

Wenhai Wang, Zhe Chen, Xiaokang Chen +8

Large language models (LLMs) have notably accelerated progress towards artificial general intelligence (AGI), with their impressive zero-shot capacity for user-tailored tasks, endo…

cs.CV20232 cited

Video Dehazing via a Multi-Range Temporal Alignment Network with Physical Prior

Jiaqi Xu, Xiaowei Hu, Lei Zhu +4

Video dehazing aims to recover haze-free frames with high visibility and contrast. This paper presents a novel framework to effectively explore the physical haze priors and aggrega…

cs.CV2023

FeatAug-DETR: Enriching One-to-Many Matching for DETRs with Feature Augmentation

Rongyao Fang, Peng Gao, Aojun Zhou +4

One-to-one matching is a crucial design in DETR-like object detection frameworks. It enables the DETR to perform end-to-end detection. However, it also faces challenges of lacking…