6 papers
Linguistic Relative Policy Optimization for Video Anomaly Reasoning
Jiaxu Leng, Jiankang Zheng, Mengjingcheng Mo +4
Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed…
SplatFont3D: Structure-Aware Text-to-3D Artistic Font Generation with Part-Level Style Control
Ji Gan, Lingxu Chen, Jiaxu Leng +1
Artistic font generation (AFG) can assist human designers in creating innovative artistic fonts. However, most previous studies primarily focus on 2D artistic fonts in flat design,…
A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
Mengjingcheng Mo, Xinyang Tong, Mingpi Tan +7
While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex…
EHGCN: Hierarchical Euclidean-Hyperbolic Fusion via Motion-Aware GCN for Hybrid Event Stream Perception
Haosheng Chen, Lian Luo, Mengjingcheng Mo +5
Event cameras, characterized by microsecond temporal resolution and very High Dynamic Range (HDR), emit high-speed event streams for perception tasks. In recent advancements, Graph…
Shape-centered Representation Learning for Visible-Infrared Person Re-identification
Shuang Li, Jiaxu Leng, Ji Gan +2
Visible-Infrared Person Re-Identification (VI-ReID) plays a critical role in all-day surveillance systems. However, existing methods primarily focus on learning appearance features…
PiercingEye: Dual-Space Video Violence Detection with Hyperbolic Vision-Language Guidance
Jiaxu Leng, Zhanjie Wu, Mingpi Tan +5
Existing weakly supervised video violence detection (VVD) methods primarily rely on Euclidean representation learning, which often struggles to distinguish visually similar yet sem…