4 papers
RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures
Cheng Li, Renjun Gao, Boyi Fu
Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous full-frame video, remain st…
LC4-DViT: Land-cover Creation for Land-cover Classification with Deformable Vision Transformer
Kai Wang, Siyi Chen, Weicong Pang +7
Land-cover underpins ecosystem services, hydrologic regulation, disaster-risk reduction, and evidence-based land planning; timely, accurate land-cover maps are therefore critical f…
MARS: Multi-Agent Robotic System with Multimodal Large Language Models for Assistive Intelligence
Renjun Gao
Multimodal large language models (MLLMs) have shown remarkable capabilities in cross-modal understanding and reasoning, offering new opportunities for intelligent assistive systems…
MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
Siyi Chen, Kai Wang, Weicong Pang +7
Land-cover understanding in remote sensing increasingly demands class-agnostic systems that generalize across datasets while remaining spatially precise and interpretable. We study…