activity
20242026
collaborators

13 papers

cs.CV2026

P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

Yi Shi, Huichao Xie, Yuqing Wang +7

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundam…

cs.CV2026

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

Haibin Tian, Huichao Xie, Xuelin Qian +3

Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing re…

cs.CV2026

LangSurf: Language-Embedded Surface Gaussians for 3D Scene Understanding

Hao Li, Minghan Qin, Zhengyu Zou +6

Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from n…

cs.CV2025

Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning

Long Li, Shuichen Ji, Ziyang Luo +4

Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify ke…

cs.CV2025

IrisNet: Infrared Image Status Awareness Meta Decoder for Infrared Small Targets Detection

Xuelin Qian, Jiaming Lu, Zixuan Wang +4

Infrared Small Target Detection (IRSTD) faces significant challenges due to low signal-to-noise ratios, complex backgrounds, and the absence of discernible target features. While d…

cs.CV2025

Dual-Granularity Semantic Prompting for Language Guidance Infrared Small Target Detection

Zixuan Wang, Haoran Sun, Jiaming Lu +5

Infrared small target detection remains challenging due to limited feature representation and severe background interference, resulting in sub-optimal performance. While recent CLI…