3 papers
cs.CV2026
SCALPEL: Semantic Cross-modal Alignment via LLM-Powered Encoder Learning for Medical Vision-Language Representation
Yunzhan Fu, Enyu Bao, Xiangyu Shen +4
Vision-language pre-training (VLP) serves as a cornerstone for medical multimodal representation learning. However, existing medical VLP frameworks are often constrained by the lim…
cs.CV2026
Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
Yihao Wu, Chenyi Xu, Liqi Yan +6
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multim…
cs.CV2025
DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery
Jiadong Tang, Yu Gao, Dianyi Yang +3
Drones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rende…