5 papers
PixelSR: Efficient Screen Content Super-Resolution via Pixel Classification
Zhiheng Li, Lei Chen, Jie Zhou +1
Screen content images are generally composed of texts and graphics. Compared to natural images, these man-made images contain a large quantity of sharp but repetitive structures. H…
FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
Yixuan Zhu, Haolin Wang, Shilin Ma +4
Recent advancements in diffusion frameworks have significantly enhanced video editing, achieving high fidelity and strong alignment with textual prompts. However, conventional appr…
InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
Yixuan Zhu, Haolin Wang, Ao Li +6
Image enhancement finds wide-ranging applications in real-world scenarios due to complex environments and the inherent limitations of imaging devices. Recent diffusion-based method…
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
Jiajun Dong, Chengkun Wang, Wenzhao Zheng +3
Effective image tokenization is crucial for both multi-modal understanding and generation tasks due to the necessity of the alignment with discrete text data. To this end, existing…
Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting
Sujia Wang, Xiangwei Shen, Yansong Tang +3
Repetitive action counting (RAC) aims to estimate the number of class-agnostic action occurrences in a video without exemplars. Most current RAC methods rely on a raw frame-to-fram…