Publications (128)
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
Zhicheng Zhao, Changfu Zhou, Yu Zhang +3
Remote Sensing Visual Question Answering (RSVQA) has gained significant research interest. However, current RSVQA methods are limited by the imaging mechanisms of optical sensors,…
Apo2Mol: 3D Molecule Generation via Dynamic Pocket-Aware Diffusion Models
Xinzhe Zheng, Shiyu Jiang, Gustavo Seabra +2
Deep generative models are rapidly advancing structure-based drug design, offering substantial promise for generating small molecule ligands that bind to specific protein targets.…
Erasure-based Interaction Network for RGBT Video Object Detection and A Unified Benchmark
Zhengzheng Tu, Qishun Wang, Hongshun Wang +2
Recently, many breakthroughs are made in the field of Video Object Detection (VOD), but the performance is still limited due to the imaging limitations of RGB sensors in adverse il…
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
Wentao Wu, Fanghua Hong, Xiao Wang +2
Existing vehicle detectors are usually obtained by training a typical detector (e.g., YOLO, RCNN, DETR series) on vehicle images based on a pre-trained backbone (e.g., ResNet, ViT)…
Alignment-Free RGBT Salient Object Detection: Semantics-guided Asymmetric Correlation Network and A Unified Benchmark
Kunpeng Wang, Danying Lin, Chenglong Li +2
RGB and Thermal (RGBT) Salient Object Detection (SOD) aims to achieve high-quality saliency prediction by exploiting the complementary information of visible and thermal image pair…
Soft Multipath Information-Based UWB Tracking in Cluttered Scenarios: Preliminaries and Validations
Chenglong Li, Zukun Lu, Long Huang +4
In this paper, we investigate ultra-wideband (UWB) localization and tracking in cluttered environments. Instead of mitigating the multipath, we exploit the specular reflections to…
CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product
Kaiwen Xue, Chenglong Li, Zhonghong Ou +10
Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. T…
FANet: Quality-Aware Feature Aggregation Network for Robust RGB-T Tracking
Yabin Zhu, Chenglong Li, Bin Luo +1
This paper investigates how to perform robust visual tracking in adverse and challenging conditions using complementary visual and thermal infrared data (RGBT tracking). We propose…
Towards Robust Optical-SAR Object Detection under Missing Modalities: A Dynamic Quality-Aware Fusion Framework
Zhicheng Zhao, Yuancheng Xu, Andong Lu +2
Optical and Synthetic Aperture Radar (SAR) fusion-based object detection has attracted significant research interest in remote sensing, as these modalities provide complementary in…
Multi-spectral Vehicle Re-identification with Cross-directional Consistency Network and a High-quality Benchmark
Aihua Zheng, Xianpeng Zhu, Zhiqi Ma +3
To tackle the challenge of vehicle re-identification (Re-ID) in complex lighting environments and diverse scenes, multi-spectral sources like visible and infrared information are t…
Graph-based Semantic Calibration Network for Unaligned UAV RGBT Image Semantic Segmentation and A Large-scale Benchmark
Fangqiang Fan, Zhicheng Zhao, Xiaoliang Ma +2
Fine-grained RGBT image semantic segmentation is crucial for all-weather unmanned aerial vehicle (UAV) scene understanding. However, UAV RGBT image semantic segmentation faces two…
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion
Xiao Wang, Jiandong Jin, Chenglong Li +3
Existing pedestrian attribute recognition (PAR) algorithms adopt pre-trained CNN (e.g., ResNet) as their backbone network for visual feature learning, which might obtain sub-optima…
Quality-Aware Multimodal Saliency Detection via Deep Reinforcement Learning
Xiao Wang, Tao Sun, Rui Yang +3
Incorporating various modes of information into the machine learning procedure is becoming a new trend. And data from various source can provide more information than single one no…
GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction
Yan Song, Zhihao Li, Chenglong Li +3
Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing methods rely on surrogate signals…
UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification
Xixi Wan, Aihua Zheng, Bo Jiang +3
Multi-modal object Re-IDentification (ReID) has gained considerable attention with the goal of retrieving specific targets across cameras using heterogeneous visual data sources. A…
Learning Where to Focus: Density-Driven Guidance for Detecting Dense Tiny Objects
Zhicheng Zhao, Xuanang Fan, Lingma Sun +2
High-resolution remote sensing imagery increasingly contains dense clusters of tiny objects, the detection of which is extremely challenging due to severe mutual occlusion and limi…
Cross-Modal UAV Object Tracking: State-Aware Representation Learning and A Unified Benchmark
Yun Xiao, Zhihong Hong, Jiandong Jin +3
Unmanned Aerial Vehicle (UAV) object tracking has emerged as a popular research field with broad practical applications. Modern UAVs are increasingly equipped with both visible lig…
Frequency-domain Multi-modal Fusion for Language-guided Medical Image Segmentation
Bo Yu, Jianhua Yang, Zetao Du +3
Automatically segmenting infected areas in radiological images is essential for diagnosing pulmonary infectious diseases. Recent studies have demonstrated that the accuracy of the…
DecoyDB: A Dataset for Graph Contrastive Learning in Protein-Ligand Binding Affinity Prediction
Yupu Zhang, Zelin Xu, Tingsong Xiao +4
Predicting the binding affinity of protein-ligand complexes plays a vital role in drug discovery. Unfortunately, progress has been hindered by the lack of large-scale and high-qual…
Prototype-based Cross-Modal Object Tracking
Lei Liu, Chenglong Li, Futian Wang +2
Cross-modal object tracking is an important research topic in the field of information fusion, and it aims to address imaging limitations in challenging scenarios by integrating sw…
Tunneling spectra of junctions for van der Waals superconductors
Yixuan Niu, Jun Cheng, Shiji Ding +5
Tunneling spectroscopy and its evolution are crucial for elucidating the intricate electronic structure and emergent phenomena in quantum materials.Nevertheless, high-quality measu…
SequencePAR: Understanding Pedestrian Attributes via A Sequence Generation Paradigm
Jiandong Jin, Xiao Wang, Yin Lin +4
Current pedestrian attribute recognition (PAR) algorithms use multi-label or multi-task learning frameworks with specific classification heads. These models often struggle with imb…
Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark
Yifei Deng, Chenglong Li, Yuyang Zhang +2
Text-aerial person retrieval aims to identify targets in UAV-captured images from eyewitness descriptions, supporting intelligent transportation and public security applications. C…
Breaking Modality Gap in RGBT Tracking: Coupled Knowledge Distillation
Andong Lu, Jiacong Zhao, Chenglong Li +2
Modality gap between RGB and thermal infrared (TIR) images is a crucial issue but often overlooked in existing RGBT tracking methods. It can be observed that modality gap mainly li…
Vehicle-centric Perception via Multimodal Structured Pre-training
Wentao Wu, Xiao Wang, Chenglong Li +2
Vehicle-centric perception plays a crucial role in many intelligent systems, including large-scale surveillance systems, intelligent transportation, and autonomous driving. Existin…
Robo-GS: A Physics Consistent Spatial-Temporal Model for Robotic Arm with Hybrid Representation
Haozhe Lou, Yurong Liu, Yike Pan +9
Real2Sim2Real plays a critical role in robotic arm control and reinforcement learning, yet bridging this gap remains a significant challenge due to the complex physical properties…
Dynamic Enhancement Network for Partial Multi-modality Person Re-identification
Aihua Zheng, Ziling He, Zi Wang +2
Many existing multi-modality studies are based on the assumption of modality integrity. However, the problem of missing arbitrary modalities is very common in real life, and this p…
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
Wentao Wu, Xiao Wang, Chenglong Li +4
Event cameras have attracted increasing attention in recent years due to their advantages in high dynamic range, high temporal resolution, low power consumption, and low latency. S…
DRGBT-1K: A Large-scale High-quality Benchmark for Dynamic RGBT Tracking
Zhaodong Ding, Chenglong Li, Zeyu Ding +2
Dynamic RGBT (DRGBT) tracking aims to continuously localize a target when the available sensing modalities and observation platforms vary over time. Compared with conventional RGBT…
Segment Any Vehicle: Semantic and Visual Context Driven SAM and A Benchmark
Xiao Wang, Ziwen Wang, Wentao Wu +4
With the rapid advancement of autonomous driving, vehicle perception, particularly detection and segmentation, has placed increasingly higher demands on algorithmic performance. Pr…
Structural Information Guided Multimodal Pre-training for Vehicle-centric Perception
Xiao Wang, Wentao Wu, Chenglong Li +4
Understanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-trai…
Disentangled Generation Network for Enlarged License Plate Recognition and A Unified Dataset
Chenglong Li, Xiaobin Yang, Guohao Wang +4
License plate recognition plays a critical role in many practical applications, but license plates of large vehicles are difficult to be recognized due to the factors of low resolu…
Learning-Driven Channel Representation for Wireless Localization: From Channel Observations to Location Inference
Hongyu Xie, Chenglong Li, Xinming Huang +4
The paper surveys learning‑driven approaches that transform wireless channel observations into representations for inferring device locations, describing frameworks, feature extrac…
AFter: Attention-based Fusion Router for RGBT Tracking
Andong Lu, Wanyu Wang, Chenglong Li +2
Multi-modal feature fusion as a core investigative component of RGBT tracking emerges numerous fusion studies in recent years. However, existing RGBT tracking methods widely adopt…
Multi-Adapter RGBT Tracking
Chenglong Li, Andong Lu, Aihua Zheng +2
The task of RGBT tracking aims to take the complementary advantages from visible spectrum and thermal infrared data to achieve robust visual tracking, and receives more and more at…
Structure and Progress Aware Diffusion for Medical Image Segmentation
Siyuan Song, Guyue Hu, Chenglong Li +3
Medical image segmentation is crucial for computer-aided diagnosis, which necessitates understanding both coarse morphological and semantic structures, as well as carving fine boun…
Cross-Modal Object Tracking: Modality-Aware Representations and A Unified Benchmark
Chenglong Li, Tianhao Zhu, Lei Liu +3
In many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in low-light conditions, and tracking performance is thus affected sig…
DeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied Tracking
Guyue Hu, Haoming Liu, Siyuan Song +3
Aerial object tracking has broad applications in public safety, emergency rescue, wildlife monitoring, and related fields. However, existing aerial tracking benchmarks are mainly b…
Semantics Guided Disentangled GAN for Chest X-ray Image Rib Segmentation
Lili Huang, Dexin Ma, Xiaowei Zhao +4
The label annotations for chest X-ray image rib segmentation are time consuming and laborious, and the labeling quality heavily relies on medical knowledge of annotators. To reduce…
Challenge-Aware RGBT Tracking
Chenglong Li, Lei Liu, Andong Lu +2
RGB and thermal source data suffer from both shared and specific challenges, and how to explore and exploit them plays a critical role to represent the target appearance in RGBT tr…
Describe and Attend to Track: Learning Natural Language guided Structural Representation and Visual Attention for Object Tracking
Xiao Wang, Chenglong Li, Rui Yang +3
The tracking-by-detection framework requires a set of positive and negative training samples to learn robust tracking models for precise localization of target objects. However, ex…
Ubiquitous Indoor Fine-Grained Positioning and Tracking: A Channel Response Perspective
Chenglong Li, Emmeric Tanghe, Sofie Pollin +1
The future of location-aided applications is shaped by the ubiquity of Internet-of-Things devices. As an increasing amount of commercial off-the-shelf radio devices support channel…
PISA: Pixelwise Image Saliency by Aggregating Complementary Appearance Contrast Measures with Edge-Preserving Coherence
Keze Wang, Liang Lin, Jiangbo Lu +2
Driven by recent vision and graphics applications such as image segmentation and object recognition, computing pixel-accurate saliency values to uniformly highlight foreground obje…
Illumination Distillation Framework for Nighttime Person Re-Identification and A New Benchmark
Andong Lu, Zhang Zhang, Yan Huang +4
Nighttime person Re-ID (person re-identification in the nighttime) is a very important and challenging task for visual surveillance but it has not been thoroughly investigated. Und…
Compositional-Degradation UAV Image Restoration: Conditional Decoupled MoE Network and A Benchmark
Jinquan Yan, Zhicheng Zhao, Zhengzheng Tu +3
UAV images are critical for applications such as large-area mapping, infrastructure inspection, and emergency response. However, in real-world flight environments, a single image i…
RGBT Tracking via Multi-Adapter Network with Hierarchical Divergence Loss
Andong Lu, Chenglong Li, Yuqing Yan +2
RGBT tracking has attracted increasing attention since RGB and thermal infrared data have strong complementary advantages, which could make trackers all-day and all-weather work. H…
Visual Tracking via Dynamic Graph Learning
Chenglong Li, Liang Lin, Wangmeng Zuo +2
Existing visual tracking methods usually localize a target object with a bounding box, in which the performance of the foreground object trackers or detectors is often affected by…
Analogies between Transformer Layers and Power Method
Chenglong Li, Claudio Altafini
In the paper we show that there is an analogy between the operations occurring in a layer of a transformer (projections and layer normalizations, disregarding the feedforward neura…
Parallel Augmentation and Dual Enhancement for Occluded Person Re-identification
Zi Wang, Huaibo Huang, Aihua Zheng +2
Occluded person re-identification (Re-ID), the task of searching for the same person's images in occluded environments, has attracted lots of attention in the past decades. Recent…
Multimodal Spatio-temporal Graph Learning for Alignment-free RGBT Video Object Detection
Qishun Wang, Zhengzheng Tu, Chenglong Li +1
RGB-Thermal Video Object Detection (RGBT VOD) can address the limitation of traditional RGB-based VOD in challenging lighting conditions, making it more practical and effective in…
Guidance Disentanglement Network for Optics-Guided Thermal UAV Image Super-Resolution
Zhicheng Zhao, Juanjuan Gu, Chenglong Li +3
Optics-guided Thermal UAV image Super-Resolution (OTUAV-SR) has attracted significant research interest due to its potential applications in security inspection, agricultural measu…
Morphological Profiling for Drug Discovery in the Era of Deep Learning
Qiaosi Tang, Ranjala Ratnayake, Gustavo Seabra +9
Morphological profiling is a valuable tool in phenotypic drug discovery. The advent of high-throughput automated imaging has enabled the capturing of a wide range of morphological…
Mixture of Enhanced-View Experts for Multi-Query Vehicle ReID and A Large-Scale Benchmark
Aihua Zheng, Jie Zhen, Chenglong Li +2
Multi-query vehicle ReID aims to leverage complementary information from diverse views for robust feature learning. However, current methods suffer from simplistic feature fusion a…
Nighttime Person Re-Identification via Collaborative Enhancement Network with Multi-domain Learning
Andong Lu, Chenglong Li, Tianrui Zha +3
Prevalent nighttime person re-identification (ReID) methods typically combine image relighting and ReID networks in a sequential manner. However, their performance (recognition acc…
Decoupled Cross-Modal Alignment Network for Text-RGBT Person Retrieval and A High-Quality Benchmark
Yifei Deng, Chenglong Li, Zhenyu Chen +2
The performance of traditional text-image person retrieval task is easily affected by lighting variations due to imaging limitations of visible spectrum sensors. In recent years, c…
Cross-modulated Attention Transformer for RGBT Tracking
Yun Xiao, Jiacong Zhao, Andong Lu +4
Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-mod…
Towards Fine-Grained Indoor Localization based on Massive MIMO-OFDM System: Experiment and Analysis
Chenglong Li, Sibren De Bast, Emmeric Tanghe +2
Fine-grained indoor localization has attracted attention recently because of the rapidly growing demand for indoor location-based services (ILBS). Specifically, massive (large-scal…
Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting
Lingbo Liu, Jiaqi Chen, Hefeng Wu +3
Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited…
Parameter-Dynamic Adaptive Fusion and Calibration Network for RGBT Tracking
Zhaoding Ding, Chenglong Li, Jiandong Jin +2
Existing RGBT trackers typically employ fusion functions with fixed parameters across different targets and scenarios. Although dynamic-architecture methods improve fusion flexibil…
Cross-Modal Object Tracking via Modality-Aware Fusion Network and A Large-Scale Dataset
Lei Liu, Mengya Zhang, Cheng Li +2
Visual tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating add…
Multi-Static UWB Radar-based Passive Human Tracking Using COTS Devices
Chenglong Li, Emmeric Tanghe, Jaron Fontaine +5
Due to its high delay resolution, the ultra-wideband (UWB) technique has been widely adopted for fine-grained indoor localization. Instead of active positioning, UWB radar-based pa…
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
Tingyu Wu, Zhisheng Chen, Ziyan Weng +8
Existing long-horizon memory benchmarks mostly use multi-turn dialogues or synthetic user histories, which makes retrieval performance an imperfect proxy for person understanding.…
RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images
Guyue Hu, Hao Song, Yuxing Tong +5
Referring detection refers to locate the target referred by natural languages, which has recently attracted growing research interests. However, existing datasets are limited to gr…
Text Analytics for Resilience-Enabled Extreme Events Reconnaissance
Alicia Y. Tsai, Selim Gunay, Minjune Hwang +4
Post-hazard reconnaissance for natural disasters (e.g., earthquakes) is important for understanding the performance of the built environment, speeding up the recovery, enhancing re…
Multi-query Vehicle Re-identification: Viewpoint-conditioned Network, Unified Dataset and New Metric
Aihua Zheng, Chaobin Zhang, Weijun Zhang +4
Existing vehicle re-identification methods mainly rely on the single query, which has limited information for vehicle representation and thus significantly hinders the performance…
Group Multi-View Transformer for 3D Shape Analysis with Spatial Encoding
Lixiang Xu, Qingzhe Cui, Richang Hong +5
In recent years, the results of view-based 3D shape recognition methods have saturated, and models with excellent performance cannot be deployed on memory-limited devices due to th…
Learning Target-oriented Dual Attention for Robust RGB-T Tracking
Rui Yang, Yabin Zhu, Xiao Wang +2
RGB-Thermal object tracking attempt to locate target object using complementary visual and thermal infrared data. Existing RGB-T trackers fuse different modalities by robust featur…
Breaking Shallow Limits: Task-Driven Pixel Fusion for Gap-free RGBT Tracking
Andong Lu, Yuanzhi Guo, Wanyu Wang +3
Current RGBT tracking methods often overlook the impact of fusion location on mitigating modality gap, which is key factor to effective tracking. Our analysis reveals that shallowe…
Tiny Object Tracking: A Large-scale Dataset and A Baseline
Yabin Zhu, Chenglong Li, Yao Liu +4
Tiny objects, frequently appearing in practical applications, have weak appearance and features, and receive increasing interests in meany vision tasks, such as object detection an…
Learning Adaptive Fusion Bank for Multi-modal Salient Object Detection
Kunpeng Wang, Zhengzheng Tu, Chenglong Li +2
Multi-modal salient object detection (MSOD) aims to boost saliency detection performance by integrating visible sources with depth or thermal infrared ones. Existing methods genera…
Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
Zhicheng Zhao, Yin Huang, Lingma Sun +2
Tiny object detection in remote sensing imagery has attracted significant research interest in recent years. Despite recent progress, achieving balanced detection performance acros…
Inference-to-complete: A High-performance and Programmable Data-plane Co-processor for Neural-network-driven Traffic Analysis
Dong Wen, Zhongpei Liu, Tong Yang +5
Neural-networks-driven intelligent data-plane (NN-driven IDP) is becoming an emerging topic for excellent accuracy and high performance. Meanwhile we argue that NN-driven IDP shoul…
Towards General Multimodal Visual Tracking
Andong Lu, Mai Wen, Jinhu Wang +4
Existing multimodal tracking studies focus on bi-modal scenarios such as RGB-Thermal, RGB-Event, and RGB-Language. Although promising tracking performance is achieved through lever…
Influence of Small Molecule Property on Antibody Response
Kai Wen, Yuchen Bai, Yujie Wei +4
Antibodies with high titer and affinity to small molecule are critical in the field for the development of vaccines against drugs of abuse, antidotes to toxins and immunoassays for…
Alignment-Free RGB-T Salient Object Detection: A Large-scale Dataset and Progressive Correlation Network
Kunpeng Wang, Keke Chen, Chenglong Li +2
Alignment-free RGB-Thermal (RGB-T) salient object detection (SOD) aims to achieve robust performance in complex scenes by directly leveraging the complementary information from una…
Octopus: A Heterogeneous In-network Computing Accelerator Enabling Deep Learning for network
Dong Wen, Tao Li, Chenglong Li +3
Deep learning (DL) for network models have achieved excellent performance in the field and are becoming a promising component in future intelligent network system. Programmable in-…
LSOTB-TIR:A Large-Scale High-Diversity Thermal Infrared Object Tracking Benchmark
Qiao Liu, Xin Li, Zhenyu He +8
In this paper, we present a Large-Scale and high-diversity general Thermal InfraRed (TIR) Object Tracking Benchmark, called LSOTBTIR, which consists of an evaluation dataset and a…
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
Yiqi Nie, Fei Wang, Junjie Chen +5
Memes represent a tightly coupled, multimodal form of social expression, in which visual context and overlaid text jointly convey nuanced affect and commentary. Inspired by cogniti…
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
Yu Zhang, Zhicheng Zhao, Ze Luo +2
Traffic scene understanding from unmanned aerial vehicle (UAV) platforms is crucial for intelligent transportation systems due to its flexible deployment and wide-area monitoring c…
RGB-T Image Saliency Detection via Collaborative Graph Learning
Zhengzheng Tu, Tian Xia, Chenglong Li +3
Image saliency detection is an active research topic in the community of computer vision and multimedia. Fusing complementary RGB and thermal infrared data has been proven to be ef…
Physics-Constrained Cross-Resolution Enhancement Network for Optics-Guided Thermal UAV Image Super-Resolution
Zhicheng Zhao, Fengjiao Peng, Jinquan Yan +3
Optics-guided thermal UAV image super-resolution has attracted significant research interest due to its potential in all-weather monitoring applications. However, existing methods…
Pedestrian Attribute Recognition: A New Benchmark Dataset and A Large Language Model Augmented Framework
Jiandong Jin, Xiao Wang, Qian Zhu +2
Pedestrian Attribute Recognition (PAR) is one of the indispensable tasks in human-centered research. However, existing datasets neglect different domains (e.g., environments, times…
Segmenting Objects in Day and Night:Edge-Conditioned CNN for Thermal Image Semantic Segmentation
Chenglong Li, Wei Xia, Yan Yan +2
Despite much research progress in image semantic segmentation, it remains challenging under adverse environmental conditions caused by imaging limitations of visible spectrum. Whil…
Hand Hygiene Assessment via Joint Step Segmentation and Key Action Scorer
Chenglong Li, Qiwen Zhu, Tubiao Liu +2
Hand hygiene is a standard six-step hand-washing action proposed by the World Health Organization (WHO). However, there is no good way to supervise medical staff to do hand hygiene…
Sparse-Dense Mixture of Experts Adapter for Multi-Modal Tracking
Yabin Zhu, Jianqi Li, Chenglong Li +3
Parameter-efficient fine-tuning (PEFT) techniques, such as prompts and adapters, are widely used in multi-modal tracking because they alleviate issues of full-model fine-tuning, in…
Dense Feature Aggregation and Pruning for RGBT Tracking
Yabin Zhu, Chenglong Li, Bin Luo +2
How to perform effective information fusion of different modalities is a core factor in boosting the performance of RGBT tracking. This paper presents a novel deep fusion algorithm…
Visible-Thermal Multiple Object Tracking: Large-scale Video Dataset and Progressive Fusion Approach
Yabin Zhu, Qianwu Wang, Chenglong Li +2
The complementary benefits from visible and thermal infrared data are widely utilized in various computer vision task, such as visual tracking, semantic segmentation and object det…
Collaborative Enhancement Network for Low-quality Multi-spectral Vehicle Re-identification
Aihua Zheng, Yongqi Sun, Zi Wang +2
The performance of multi-spectral vehicle Re-identification (ReID) is significantly degraded when some important discriminative cues in visible, near infrared and thermal infrared…
LasHeR: A Large-scale High-diversity Benchmark for RGBT Tracking
Chenglong Li, Wanlin Xue, Yaqing Jia +4
RGBT tracking receives a surge of interest in the computer vision community, but this research field lacks a large-scale and high-diversity benchmark dataset, which is essential fo…
RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
Andong Lu, Wanyu Wang, Chenglong Li +2
Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, w…
Adapting Segment Anything Model to Multi-modal Salient Object Detection with Semantic Feature Fusion Guidance
Kunpeng Wang, Danying Lin, Chenglong Li +2
Although most existing multi-modal salient object detection (SOD) methods demonstrate effectiveness through training models from scratch, the limited multi-modal data hinders these…
RGBT Tracking via Progressive Fusion Transformer with Dynamically Guided Learning
Yabin Zhu, Chenglong Li, Xiao Wang +2
Existing Transformer-based RGBT tracking methods either use cross-attention to fuse the two modalities, or use self-attention and cross-attention to model both modality-specific an…
Modality-missing RGBT Tracking: Invertible Prompt Learning and High-quality Benchmarks
Andong Lu, Jiacong Zhao, Chenglong Li +2
Current RGBT tracking research relies on the complete multi-modal input, but modal information might miss due to some factors such as thermal sensor self-calibration and data trans…
An Empirical Study of Mamba-based Pedestrian Attribute Recognition
Xiao Wang, Weizhe Kong, Jiandong Jin +5
Current strong pedestrian attribute recognition models are developed based on Transformer networks, which are computationally heavy. Recently proposed models with linear complexity…
Attributes Guided Feature Learning for Vehicle Re-identification
Hongchao Li, Xianmin Lin, Aihua Zheng +4
Vehicle Re-ID has recently attracted enthusiastic attention due to its potential applications in smart city and urban surveillance. However, it suffers from large intra-class varia…
Design of an electron gun for terahertz radiation source
Ji Li, Y. J. Pei, Tongning Hu +4
With the aim to obtain short-pulse bunches with high peak current for a terahertz radiation source, an EC-ITC (External-Cathode Independently Tunable Cells) RF gun was employed. As…
Phase-based Variant Maximum Likelihood Positioning for Passive UHF-RFID Tags
Chenglong Li, Emmeric Tanghe, David Plets +6
Radio frequency identification (RFID) technology brings tremendous advancement in Internet-of-Things, especially in supply chain and smart inventory management. Phase-based passive…
RGBT Salient Object Detection: A Large-scale Dataset and Benchmark
Zhengzheng Tu, Yan Ma, Zhun Li +3
Salient object detection in complex scenes and environments is a challenging research topic. Most works focus on RGB-based salient object detection, which limits its performance of…
Dynamic Disentangled Fusion Network for RGBT Tracking
Chenglong Li, Tao Wang, Zhaodong Ding +2
RGBT tracking usually suffers from various challenging factors of low resolution, similar appearance, extreme illumination, thermal crossover and occlusion, to name a few. Existing…
DehazeMamba: SAR-guided Optical Remote Sensing Image Dehazing with Adaptive State Space Model
Zhicheng Zhao, Jinquan Yan, Chenglong Li +2
Optical remote sensing image dehazing presents significant challenges due to its extensive spatial scale and highly non-uniform haze distribution, which traditional single-image de…