61 citations · 70 across the 6 of their papers we have counts for
6 papers
Rethinking Domain Adaptation and Generalization in the Era of CLIP
Ruoyu Feng, Tao Yu, Xin Jin +3
In recent studies on domain adaptation, significant emphasis has been placed on the advancement of learning shared knowledge from a source domain to a target domain. Recently, the…
MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
Yanhui Wang, Jianmin Bao, Wenming Weng +12
We present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with vi…
Prompt-ICM: A Unified Framework towards Image Coding for Machines with Task-driven Prompts
Ruoyu Feng, Jinming Liu, Xin Jin +3
Image coding for machines (ICM) aims to compress images to support downstream AI analysis instead of human perception. For ICM, developing a unified codec to reduce information red…
Inpaint Anything: Segment Anything Meets Image Inpainting
Tao Yu, Runseng Feng, Ruoyu Feng +4
Modern image inpainting systems, despite the significant progress, often struggle with mask selection and holes filling. Based on Segment-Anything Model (SAM), we make the first at…
HST: Hierarchical Swin Transformer for Compressed Image Super-resolution
Bingchen Li, Xin Li, Yiting Lu +3
Compressed Image Super-resolution has achieved great attention in recent years, where images are degraded with compression artifacts and low-resolution artifacts. Since the complex…
Semantically Video Coding: Instill Static-Dynamic Clues into Structured Bitstream for AI Tasks
Xin Jin, Ruoyu Feng, Simeng Sun +3
Traditional media coding schemes typically encode image/video into a semantic-unknown binary stream, which fails to directly support downstream intelligent tasks at the bitstream l…