13 papers
P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation
Yi Shi, Huichao Xie, Yuqing Wang +7
Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundam…
UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery
Haibin Tian, Huichao Xie, Xuelin Qian +3
Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing re…
LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding
Hao Li, Minghan Qin, Zhengyu Zou +6
Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from n…
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
Long Li, Shuichen Ji, Ziyang Luo +4
Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify ke…
IrisNet: Infrared Image Status Awareness Meta Decoder for Infrared Small Targets Detection
Xuelin Qian, Jiaming Lu, Zixuan Wang +4
Infrared Small Target Detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, complex backgrounds, and the absence of discernible target features. While d…
Dual-Granularity Semantic Prompting for Language Guidance Infrared Small Target Detection
Zixuan Wang, Haoran Sun, Jiaming Lu +5
Infrared small target detection remains challenging due to limited feature representation and severe background interference, resulting in sub-optimal performance. While recent CLI…