activity
20222025
collaborators

8 papers

cs.CV2025

Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation

Shihao Zhao, Yitong Chen, Zeyinzi Jiang +5

Unified understanding and generation is a highly appealing research direction in multimodal learning. There exist two approaches: one trains a transformer via an auto-regressive pa…

cs.CV2025

NTIRE 2025 Challenge on UGC Video Enhancement: Methods and Results

Nikolay Safonov, Alexey Bryncev, Andrey Moskalenko +28

This paper presents an overview of the NTIRE 2025 Challenge on UGC Video Enhancement. The challenge constructed a set of 150 user-generated content videos without reference ground…

cs.CV2025

Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists

Bojia Zi, Penghui Ruan, Marco Chen +7

Recent advancements in video generation have spurred the development of video editing techniques, which can be divided into inversion-based and end-to-end methods. However, current…

cs.CV2024

Elucidating the design space of language models for image generation

Xuantong Liu, Shaozhe Hao, Xianbiao Qi +4

The success of autoregressive (AR) language models in text generation has inspired the computer vision community to adopt Large Language Models (LLMs) for image generation. However…

cs.CV2024

BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities

Shaozhe Hao, Xuantong Liu, Xianbiao Qi +5

We introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation ca…

cs.CV2024

CusConcept: Customized Visual Concept Decomposition with Diffusion Models

Zhi Xu, Shaozhe Hao, Kai Han

Enabling generative models to decompose visual concepts from a single image is a complex and challenging problem. In this paper, we study a new and challenging task, customized con…