Publications (33)
VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking
Zekun Qian, Ruize Han, Junhui Hou +2
Open-vocabulary multi-object tracking (OVMOT) represents a critical new challenge involving the detection and tracking of diverse object categories in videos, encompassing both see…
VFMSDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection
Yupeng Zhang, Ruize Han, Ningnan Guo +3
Real-world weather, illumination, and imaging variations often induce severe domain shifts, degrading single-source detectors in unseen environments. Existing single-domain general…
Online Reasoning Video Object Segmentation
Jinyuan Liu, Yang Wang, Zeyu Zhao +3
Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded references. However, existi…
From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera Calibration
Zekun Qian, Ruize Han, Wei Feng +2
We tackle a new problem of multi-view camera and subject registration in the bird's eye view (BEV) without pre-given camera calibration. This is a very challenging problem since it…
Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference
Tong Jin, Yunpeng Liu, Shuyu Hu +4
Recent visual place recognition (VPR) methods based on vision transformers, particularly foundation models, have achieved remarkable recognition performance. However, these models…
CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking
Xiangqun Zhang, Likai Wang, Zekun Qian +2
The paper introduces CD-RMOT-Bench, a benchmark for evaluating how well referring multi-object tracking models trained on one visual domain perform on different, unseen domains, an…
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
Yupeng Zhang, Ruize Han, Fangnan Zhou +2
Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of shifts often occur simultaneously…
From Indoor To Outdoor: Unsupervised Domain Adaptive Gait Recognition
Likai Wang, Ruize Han, Wei Feng +1
Gait recognition is an important AI task, which has been progressed rapidly with the development of deep learning. However, existing learning based gait recognition methods mainly…
CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
Likai Wang, Ruize Han, Xiangqun Zhang +1
Vehicles, as one of the most common and significant objects in the real world, the researches on which using computer vision technologies have made remarkable progress, such as veh…
COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm
Zekun Qian, Wei Feng, Ruize Han +1
Multi-Object Tracking (MOT) has traditionally focused on a few specific categories, restricting its applicability to real-world scenarios involving diverse objects. Open-Vocabulary…
Combining the Silhouette and Skeleton Data for Gait Recognition
Likai Wang, Ruize Han, Wei Feng
Gait recognition, a long-distance biometric technology, has aroused intense interest recently. Currently, the two dominant gait recognition works are appearance-based and model-bas…
Multiple Human Association between Top and Horizontal Views by Matching Subjects' Spatial Distributions
Ruize Han, Yujun Zhang, Wei Feng +5
Video surveillance can be significantly enhanced by using both top-view data, e.g., those from drone-mounted cameras in the air, and horizontal-view data, e.g., those from wearable…
Single Object Tracking Research: A Survey
Ruize Han, Wei Feng, Qing Guo +1
Visual object tracking is an important task in computer vision, which has many real-world applications, e.g., video surveillance, visual navigation. Visual object tracking also has…
The Fourth Challenge on Image Super-Resolution (4) at NTIRE 2026: Benchmark Results and Method Overview
Zheng Chen, Kai Liu, Jingkai Wang +150
This paper presents the NTIRE 2026 image super-resolution (4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to r…
LV-OSD: Language-Vision-Complementary Open-Set Object Detection
Yupeng Zhang, Ruize Han, Wei Feng +2
Object detection is an important task in computer vision, which aims to detect the objects of interest. through the given category list or query images. In this work, we propose a…
BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
Zekun Qian, Ruize Han, Wei Feng
Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack exp…
OCTrack: Benchmarking the Open-Corpus Multi-Object Tracking
Zekun Qian, Ruize Han, Wei Feng +3
We study a novel yet practical problem of open-corpus multi-object tracking (OCMOT), which extends the MOT into localizing, associating, and recognizing generic-category objects of…
Robust Collaborative Perception without External Localization and Clock Devices
Zixing Lei, Zhenyang Ni, Ruize Han +5
A consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information excha…
ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification
Yupeng Zhang, Yuzhong Feng, Ruize Han +3
Open-domain open-vocabulary detection (ODOVD) requires detectors to generalize to both novel categories and unseen domains, making it more challenging than open-vocabulary detectio…
Unveiling the Power of Self-supervision for Multi-view Multi-human Association and Tracking
Wei Feng, Feifan Wang, Ruize Han +2
Multi-view multi-human association and tracking (MvMHAT), is a new but important problem for multi-person scene video surveillance, aiming to track a group of people over time in e…
Access Trends of In-network Cache for Scientific Data
Ruize Han, Alex Sim, Kesheng Wu +6
Scientific collaborations are increasingly relying on large volumes of data for their work and many of them employ tiered systems to replicate the data to their worldwide user comm…
Self-supervised Social Relation Representation for Human Group Detection
Jiacheng Li, Ruize Han, Haomin Yan +3
Human group detection, which splits crowd of people into groups, is an important step for video-based human social activity analysis. The core of human group detection is the human…
OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking
Haiji Liang, Ruize Han
Open-vocabulary object perception has become an important topic in artificial intelligence, which aims to identify objects with novel classes that have not been seen during trainin…
NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
Yupeng Zhang, Ruize Han, Zhiwei Chen +2
Despite the remarkable progress in open-vocabulary object detection (OVD), a significant gap remains between the training and testing phases. During training, the RPN and RoI heads…
Synthetic-To-Real Video Person Re-ID
Xiangqun Zhang, Wei Feng, Ruize Han +3
Person re-identification (Re-ID) is an important task and has significant applications for public security and information forensics, which has progressed rapidly with the developm…
Panoramic Human Activity Recognition
Ruize Han, Haomin Yan, Jiacheng Li +3
To obtain a more comprehensive activity understanding for a crowded scene, in this paper, we propose a new problem of panoramic human activity recognition (PAR), which aims to simu…
Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms
Ruiqi Wu, Bingliang Jiao, Ruize Han +6
Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory…
A Benchmark of Video-Based Clothes-Changing Person Re-Identification
Likai Wang, Xiangqun Zhang, Ruize Han +4
Person re-identification (Re-ID) is a classical computer vision task and has achieved great progress so far. Recently, long-term Re-ID with clothes-changing has attracted increasin…
Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models
Yifeng Wu, Huimin Huang, Ruiluo Wu +5
Multimodal Large Language Models (MLLMs) often lose track of the right image regions during fine-grained spatial reasoning, because a textual query rarely carries any explicit geom…
IMWM: Intuition Models Complement World Models for Latent Planning
Baoqi Gao, Ruize Han, Miao Wang +1
Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a p…
RT-SDGOD: Real-Time Single-Domain Generalized Object Detection
Yupeng Zhang, Fangzhuo Gao, Ruize Han +2
In real-world deployment under strict real-time constraints, weather and imaging variations induce significant distribution shifts, severely degrading detectors. Single-Domain Gene…
The First Challenge on Remote Sensing Infrared Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview
Kai Liu, Haoyang Yue, Zeli Lin +65
This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super-Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The challenge aims to recover high…
COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection
Yupeng Zhang, Ruize Han, Yuzhong Feng +3
Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-world category spaces continual…