7 papers · 1 filter
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
Long Li, Shuichen Ji, Ziyang Luo +4
Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify ke…
Hierarchical Mixing Architecture for Low-light RAW Image Enhancement
Xianmin Chen, Peiliang Huang, Longfei Han +2
With the rapid development of deep learning, low-light RAW image enhancement (LLRIE) has achieved remarkable progress. However, the challenge that how to simultaneously achieve str…
PolSAM: Polarimetric Scattering Mechanism Informed Segment Anything Model
Yuqing Wang, Zhongling Huang, Shuxin Yang +4
PolSAR data presents unique challenges due to its rich and complex characteristics. Existing data representations, such as complex-valued data, polarimetric features, and amplitude…
Retinex-RAWMamba: Bridging Demosaicing and Denoising for Low-Light RAW Image Enhancement
Xianmin Chen, Longfei Han, Peiliang Huang +3
Low-light image enhancement, particularly in cross-domain tasks such as mapping from the raw domain to the sRGB domain, remains a significant challenge. Many deep learning-based me…
Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction
Dong Li, Wenqi Zhong, Wei Yu +5
Video virtual try-on aims to seamlessly dress a subject in a video with a specific garment. The primary challenge involves preserving the visual authenticity of the garment while d…
DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
Hao Li, Yuanyuan Gao, Haosong Peng +7
Novel-view synthesis (NVS) approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, mak…