9 papers
Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models
Liangsheng Liu, Si Chen, Jiamin Wu +5
Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posing serious risks in real-worl…
FS-I2P:A Hierarchical Focus-Sweep Registration Network with Dynamically Allocated Depth
Zhixin Cheng, Yujia Chen, Xujing Tao +4
Image-to-point cloud registration is often challenged by viewpoint changes, cross-modal discrepancies, and repetitive textures, which induce scale ambiguity and consequently lead t…
Sat3R: Satellite DSM Reconstruction via RPC-Aware Depth Fine-tuning
Qiaoyi Yang, Chaoyi Zhou, Xi Liu +9
Accurate Digital Surface Model (DSM) reconstruction from satellite imagery is critical for applications such as disaster response, urban planning, and large-scale geographic mappin…
GLASS: Geometry-aware Local Alignment and Structure Synchronization Network for 2D-3D Registration
Zhixin Cheng, Jiacheng Deng, Xinjun Li +5
Image-to-point cloud registration methods typically follow a coarse-to-fine pipeline, extracting patch-level correspondences and refining them into dense pixel-to-point matches. Ho…
GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
Xujing Tao, Chuxin Wang, Yubo Ai +8
Open-vocabulary 3D semantic segmentation aims to segment arbitrary categories beyond the training set. Existing methods predominantly rely on distilling knowledge from 2D open-voca…
VCR: Variance-Driven Channel Recalibration for Robust Low-Light Enhancement
Zhixin Cheng, Fangwen Zhang, Xiaotian Yin +2
Most sRGB-based LLIE methods suffer from entangled luminance and color, while the HSV color space offers insufficient decoupling at the cost of introducing significant red and blac…