7 papers
Noise Self-Regression: A New Learning Paradigm to Enhance Low-Light Images Without Task-Related Data
Zhao Zhang, Suiyi Zhao, Xiaojie Jin +4
Deep learning-based low-light image enhancement (LLIE) is a task of leveraging deep neural networks to enhance the image illumination while keeping the image content unchanged. Fro…
Two are better than one: Context window extension with multi-grained self-injection
Wei Han, Pan Zhou, Soujanya Poria +1
The limited context window of contemporary large language models (LLMs) remains a huge barrier to their broader application across various domains. While continual pre-training on…
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
Hao Fei, Shengqiong Wu, Meishan Zhang +3
While pre-training large-scale video-language models (VLMs) has shown remarkable potential for various downstream video-language tasks, existing VLMs can still suffer from certain…
Gamba: Marry Gaussian Splatting with Mamba for single view 3D reconstruction
Qiuhong Shen, Zike Wu, Xuanyu Yi +4
We tackle the challenge of efficiently reconstructing a 3D asset from a single image at millisecond speed. Existing methods for single-image 3D reconstruction are primarily based o…
UniParser: Multi-Human Parsing with Unified Correlation Representation Learning
Jiaming Chu, Lei Jin, Junliang Xing +1
Multi-human parsing is an image segmentation task necessitating both instance-level and fine-grained category-level information. However, prior research has typically processed the…
Instant3D: Instant Text-to-3D Generation
Ming Li, Pan Zhou, Jia-Wei Liu +4
Text-to-3D generation has attracted much attention from the computer vision community. Existing methods mainly optimize a neural field from scratch for each text prompt, relying on…