7 papers
AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
Jirong Zha, Yuxuan Fan, Tianyu Zhang +4
Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap…
Flight Dynamics to Sensing Modalities: Exploiting Drone Ground Effect for Accurate Edge Detection
Chenyu Zhao, Jingao Xu, Ciyu Ruan +9
Drone-based rapid and accurate environmental edge detection is highly advantageous for tasks such as disaster relief and autonomous navigation. Current methods, using radars or cam…
DIMM: Decoupled Multi-hierarchy Kalman Filter for 3D Object Tracking
Jirong Zha, Yuxuan Fan, Kai Li +4
State estimation is challenging for 3D object tracking with high maneuverability, as the target's state transition function changes rapidly, irregularly, and is unknown to the esti…
A Unified and Scalable Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
Jie Zhu, Jirong Zha, Ding Li +1
Self-supervised learning shows promise in harnessing extensive unlabeled data, but it also confronts significant privacy concerns, especially in vision. In this paper, we perform m…
Scalable UAV Multi-Hop Networking via Multi-Agent Reinforcement Learning with Large Language Models
Yanggang Xu, Jirong Zha, Weijie Hong +5
In disaster scenarios, establishing robust emergency communication networks is critical, and unmanned aerial vehicles (UAVs) offer a promising solution to rapidly restore connectiv…
How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM
Jirong Zha, Yuxuan Fan, Xiao Yang +2
3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs)…