5 papers
WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing
Ming Hu, Mingyu Dou, Jianfu Yin +5
Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing one-step editing methods pr…
UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation
Mingyu Dou, Shi Qiu, Ming Hu +2
Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video perception is largely limited to box…
Leveraging Multimodal LLMs for Built Environment and Housing Attribute Assessment from Street-View Imagery
Siyuan Yao, Siavash Ghorbany, Kuangshi Ai +6
We present a novel framework for automatically evaluating building conditions nationwide in the United States by leveraging large language models (LLMs) and Google Street View (GSV…
FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
Ming Hu, Yongsheng Huo, Mingyu Dou +6
Fine-grained anomaly detection is crucial in industrial and medical applications, but labeled anomalies are often scarce, making zero-shot detection challenging. While vision-langu…
AdaptOVCD: Training-Free Open-Vocabulary Remote Sensing Change Detection via Adaptive Information Fusion
Mingyu Dou, Shi Qiu, Ming Hu +4
Remote sensing change detection plays a pivotal role in domains such as environmental monitoring, urban planning, and disaster assessment. However, existing methods typically rely…