activity
20172023
most citedMutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identification

392 citations · 1.4k across the 54 of their papers we have counts for

collaborators

89 papers

cs.CV202347 cited

Meta-Transformer: A Unified Framework for Multimodal Learning

Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang +4

Multimodal learning aims to build models that can process and relate information from multiple modalities. Despite years of development in this field, it still remains challenging…

cs.CV2023

JourneyDB: A Benchmark for Generative Image Understanding

Keqiang Sun, Junting Pan, Yuying Ge +11

While recent advancements in vision-language models have had a transformative impact on multi-modal comprehension, the extent to which these models possess the ability to comprehen…

cs.CV2023

Retrieving-to-Answer: Zero-Shot Video Question Answering with Frozen Large Language Models

Junting Pan, Ziyi Lin, Yuying Ge +5

Video Question Answering (VideoQA) has been significantly advanced from the scaling of recent Large Language Models (LLMs). The key idea is to convert the visual information into t…

cs.CV2023

Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Xiaoshi Wu, Yiming Hao, Keqiang Sun +4

Recent text-to-image generative models can generate high-fidelity images from text inputs, but the quality of these generated images cannot be accurately evaluated by existing eval…

cs.CV20237 cited

Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

Fu-Yun Wang, Wenshuo Chen, Guanglu Song +3

Leveraging large-scale image-text datasets and advancements in diffusion models, text-driven generative models have made remarkable strides in the field of image generation and edi…

cs.RO202331 cited

Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Siyuan Huang, Zhengkai Jiang, Hao Dong +3

Foundation models have made significant strides in various applications, including text-to-image generation, panoptic segmentation, and natural language processing. This paper pres…