2 papers
cs.CV2026
Visual Instruction Pretraining for Domain-Specific Foundation Models
Yuxuan Li, Yicheng Zhang, Wenhao Tang +4
Modern computer vision is converging on a closed loop in which perception, reasoning and generation mutually reinforce each other. However, this loop remains incomplete: the top-do…
cs.CV2025
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
Yuxuan Li, Xiang Li, Yunheng Li +5
With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional Object detection models are trained on a si…