activity
20142024
most citedLearning Deep Structured Models

115 citations · 169 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CV2024

COCONut: Modernizing COCO Segmentation

Xueqing Deng, Qihang Yu, Peng Wang +2

In recent decades, the vision community has witnessed remarkable progress in visual recognition, partially owing to advancements in dataset benchmarks. Notably, the established COC…

cs.CV20241 cited

ViTamin: Designing Scalable Vision Models in the Vision-Language Era

Jieneng Chen, Qihang Yu, Xiaohui Shen +2

Recent breakthroughs in vision-language models (VLMs) start a new page in the vision community. The VLMs provide stronger and more generalizable feature embeddings compared to thos…

cs.CV20246 cited

SPFormer: Enhancing Vision Transformer with Superpixel Representation

Jieru Mei, Liang-Chieh Chen, Alan Yuille +1

In this work, we introduce SPFormer, a novel Vision Transformer enhanced by superpixel representation. Addressing the limitations of traditional Vision Transformers' fixed-size, no…

cs.CV20231 cited

Towards Open-Ended Visual Recognition with Large Language Model

Qihang Yu, Xiaohui Shen, Liang-Chieh Chen

Localizing and recognizing objects in the open-ended physical world poses a long-standing challenge within the domain of machine perception. Recent methods have endeavored to addre…

cs.CV20231 cited

PolyMaX: General Dense Prediction with Mask Transformer

Xuan Yang, Liangzhe Yuan, Kimberly Wilber +8

Dense prediction tasks, such as semantic segmentation, depth estimation, and surface normal prediction, can be easily formulated as per-pixel classification (discrete outputs) or r…

cs.CV202331 cited

Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP

Qihang Yu, Ju He, Xueqing Deng +2

Open-vocabulary segmentation is a challenging task requiring segmenting and recognizing objects from an open set of categories. One way to address this challenge is to leverage mul…