activity
20242026
collaborators

16 papers

cs.CV2026

DisDop: Distillation with Domain Priors for Open-Vocabulary Aerial Object Detection

Ruihao Xu, Yong Liu, Yansong Tang +6

With the widespread application of drones in recent years, object detection of aerial images has attracted increasing attention, especially open-vocabulary aerial detection which i…

cs.CV2026

FDDet: Achieving Data-Efficient Food Defect Detection Under Real-World Scenarios

Ruihao Xu, Yong Liu, Yansong Tang

Food defect detection is critical for automated quality control, yet existing studies lack unified benchmarks and suffer from data scarcity. We introduce FDD-48, a comprehensive da…

cs.CV2026

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis

Ruihao Xu, Xingming Shui, Jingxuan Niu +4

As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evidence and traceable accountabili…

cs.CV2026

Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

Deyi Zhu, Yuji Wang, Yong Liu +4

Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects and challenging scenarios with…

cs.CV2026

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

Yiyang Fu, Chubin Zhang, Shukai Gong +7

It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-Language-Action (VLA) models…

cs.CV2026

Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation

Sule Bai, Yong Liu, Yifei Han +4

Recent advancements in pre-trained vision-language models like CLIP have enabled the task of open-vocabulary segmentation. CLIP demonstrates impressive zero-shot capabilities in va…