5 citations · 10 across the 5 of their papers we have counts for
8 papers · 1 filter
GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?
Ruihang Li, Leigang Qu, Jingxu Zhang +6
The rapid advancement of visual generation models has outpaced traditional evaluation approaches, necessitating the adoption of Vision-Language Models as surrogate judges. In this…
Distribution Matching Variational AutoEncoder
Sen Ye, Jianning Pei, Mengde Xu +4
Most visual generative models compress images into a latent space before applying diffusion or autoregressive modelling. Yet, existing approaches such as VAEs and foundation model…
Tokenize Image as a Set
Zigang Geng, Mengde Xu, Han Hu +1
This paper proposes a fundamentally new paradigm for image generation through set-based tokenization and distribution modeling. Unlike conventional methods that serialize images in…
Side Adapter Network for Open-Vocabulary Semantic Segmentation
Mengde Xu, Zheng Zhang, Fangyun Wei +2
This paper presents a new framework for open-vocabulary semantic segmentation with the pre-trained vision-language model, named Side Adapter Network (SAN). Our approach models the…
Bootstrap Your Object Detector via Mixed Training
Mengde Xu, Zheng Zhang, Fangyun Wei +5
We introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by ut…
End-to-End Semi-Supervised Object Detection with Soft Teacher
Mengde Xu, Zheng Zhang, Han Hu +5
This paper presents an end-to-end semi-supervised object detection approach, in contrast to previous more complex multi-stage methods. The end-to-end training gradually improves ps…