activity
20232026
collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

PixelSR: Efficient Screen Content Super-Resolution via Pixel Classification

Zhiheng Li, Lei Chen, Jie Zhou +1

Screen content images are generally composed of texts and graphics. Compared to natural images, these man-made images contain a large quantity of sharp but repetitive structures. H…

cs.CV2026

DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection

Ruihao Xu, Yong Liu, Yansong Tang +6

With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocabulary aerial detection which i…

cs.CV2025

Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline

Linqing Zhao, Xiuwei Xu, Yirui Wang +5

Incrementally recovering real-sized 3D geometry from a pose-free RGB stream is a challenging task in 3D reconstruction, requiring minimal assumptions on input data. Existing method…

cs.CV2025

InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

Yixuan Zhu, Haolin Wang, Ao Li +6

Image enhancement finds wide-ranging applications in real-world scenarios due to complex environments and the inherent limitations of imaging devices. Recent diffusion-based method…

cs.CV2025

GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting

Jiajun Dong, Chengkun Wang, Wenzhao Zheng +3

Effective image tokenization is crucial for both multi-modal understanding and generation tasks due to the necessity of the alignment with discrete text data. To this end, existing…

cs.CV2024

GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian Generation

Chubin Zhang, Hongliang Song, Yi Wei +3

In this work, we introduce the Geometry-Aware Large Reconstruction Model (GeoLRM), an approach which can predict high-quality assets with 512k Gaussians and 21 input images in only…