7 papers
Vehicle-centric Perception via Multimodal Structured Pre-training
Wentao Wu, Xiao Wang, Chenglong Li +2
Vehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existin…
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
Xiao Wang, Ziwen Wang, Wentao Wu +4
With the rapid advancement of autonomous driving, vehicle perception, particularly detection and segmentation, has placed increasingly higher demands on algorithmic performance. Pr…
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
Wentao Wu, Xiao Wang, Chenglong Li +4
Event cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. S…
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
Xinnan Zhu, Yicheng Zhu, Tixin Chen +2
Temporal action detection aims to locate and classify actions in untrimmed videos. While recent works focus on designing powerful feature processors for pre-trained representations…
Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection
Wentao Wu, Chenglong Li, Xiao Wang +2
Existing multimodal UAV object detection methods often overlook the impact of semantic gaps between modalities, which makes it difficult to achieve accurate semantic and spatial al…
XiHeFusion: Harnessing Large Language Models for Science Communication in Nuclear Fusion
Xiao Wang, Qingquan Yang, Fuling Wang +12
Nuclear fusion is one of the most promising ways for humans to obtain infinite energy. Currently, with the rapid development of artificial intelligence, the mission of nuclear fusi…