most citedVision-Language Intelligence: Tasks, Representation Learning, and Large Models

30 citations · 51 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV20231 cited

Visual In-Context Prompting

Feng Li, Qing Jiang, Hao Zhang +9

In-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existin…

cs.CV202315 cited

detrex: Benchmarking Detection Transformers

Tianhe Ren, Shilong Liu, Feng Li +13

The DEtection TRansformer (DETR) algorithm has received considerable attention in the research community and is gradually emerging as a mainstream approach for object detection and…

cs.CV20223 cited

DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding

Shilong Liu, Yaoyuan Liang, Feng Li +5

In this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG req…

cs.CV20222 cited

A Unified Mutual Supervision Framework for Referring Expression Segmentation and Generation

Shijia Huang, Feng Li, Hao Zhang +3

Reference Expression Segmentation (RES) and Reference Expression Generation (REG) are mutually inverse tasks that can be naturally jointly trained. Though recent work has explored…

cs.CV202230 cited

Vision-Language Intelligence: Tasks, Representation Learning, and Large Models

Feng Li, Hao Zhang, Yi-Fan Zhang +5

This paper presents a comprehensive survey of vision-language (VL) intelligence from the perspective of time. This survey is inspired by the remarkable progress in both computer vi…