activity
20242026
collaborators

11 papers

cs.CV2026

FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion

Yunzhong Si, Huiying Xu, Xinzhong Zhu +4

Small object detection in Unmanned Aerial Vehicle (UAV) imagery remains challenging under adverse conditions, including complex weather, low illumination, and sensor noise. These c…

cs.CR2026

GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio

Zhenhao Zhu, Yue Liu, Yanpei Guo +9

We present GuardReasoner-Omni, a reasoning-based guardrail model designed to moderate text, image, video, and audio data. First, we construct a comprehensive training corpus compri…

cs.CV2026

Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection

Zhiwei Ning, Xuanang Gao, Jiaxi Cao +5

Linear modeling methods like Mamba have been merged as the effective backbone for the 3D object detection task. However, previous Mamba-based methods utilize the bidirectional enco…

cs.CR2025

ExtendAttack: Attacking Servers of LRMs via Extending Reasoning

Zhenhao Zhu, Yue Liu, Zhiwei Xu +9

Large Reasoning Models (LRMs) have demonstrated promising performance in complex tasks. However, the resource-consuming reasoning processes may be exploited by attackers to malicio…

cs.CV2025

Point Cloud Quantization through Multimodal Prompting for 3D Understanding

Hongxuan Li, Wencheng Zhu, Huiying Xu +2

Vector quantization has emerged as a powerful tool in large-scale multimodal models, unifying heterogeneous representations through discrete token encoding. However, its effectiven…

cs.LG2025

RAC-DMVC: Reliability-Aware Contrastive Deep Multi-View Clustering under Multi-Source Noise

Shihao Dong, Yue Liu, Xiaotong Zhou +3

Multi-view clustering (MVC), which aims to separate the multi-view data into distinct clusters in an unsupervised manner, is a fundamental yet challenging task. To enhance its appl…