activity
20212023
most citedTinyViT: Fast Pretraining Distillation for Small Vision Transformers

22 citations · 67 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CV20232 cited

HQ-50K: A Large-scale, High-quality Dataset for Image Restoration

Qinhong Yang, Dongdong Chen, Zhentao Tan +6

This paper introduces a new large-scale image restoration dataset, called HQ-50K, which contains 50,000 high-quality images with rich texture details and semantic diversity. We ana…

cs.CV20235 cited

Designing a Better Asymmetric VQGAN for StableDiffusion

Zixin Zhu, Xuelu Feng, Dongdong Chen +5

StableDiffusion is a revolutionary text-to-image generator that is causing a stir in the world of image generation and editing. Unlike traditional methods that learn a diffusion mo…

cs.CV20231 cited

Album Storytelling with Iterative Story-aware Captioning and Large Language Models

Munan Ning, Yujia Xie, Dongdong Chen +5

This work studies how to transform an album to vivid and coherent stories, a task we refer to as "album storytelling". While this task can help preserve memories and facilitate exp…

cs.CL20232 cited

i-Code Studio: A Configurable and Composable Framework for Integrative AI

Yuwei Fang, Mahmoud Khademi, Chenguang Zhu +8

Artificial General Intelligence (AGI) requires comprehensive understanding and generation capabilities for a variety of tasks spanning different modalities and functionalities. Int…

cs.CL20231 cited

i-Code V2: An Autoregressive Generation Framework over Vision, Language, and Speech Data

Ziyi Yang, Mahmoud Khademi, Yichong Xu +16

The convergence of text, visual, and audio data is a key step towards human-like artificial intelligence, however the current Vision-Language-Speech landscape is dominated by encod…

cs.CV20237 cited

ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Junke Wang, Dongdong Chen, Chong Luo +4

Existing deep video models are limited by specific tasks, fixed input-output spaces, and poor generalization capabilities, making it difficult to deploy them in real-world scenario…