2 papers
cs.LG2025
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Jing Liu, Sihan Chen, Xingjian He +4
In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-langu…
cs.CV2024
SLLEN: Semantic-aware Low-light Image Enhancement Network
Mingye Ju, Chuheng Chen, Charles A. Guo +3
How to effectively explore semantic feature is vital for low-light image enhancement (LLE). Existing methods usually utilize the semantic feature that is only drawn from the output…