papers

Publications (33)

cs.CV2024

VOVTrack: Exploring the Potentiality in Videos for Open-Vocabulary Object Tracking

Zekun Qian, Ruize Han, Junhui Hou +2

Open-vocabulary multi-object tracking (OVMOT) represents a critical new challenge involving the detection and tracking of diverse object categories in videos, encompassing both see…

cs.CV2026

VFMSDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection

Yupeng Zhang, Ruize Han, Ningnan Guo +3

Real-world weather, illumination, and imaging variations often induce severe domain shifts, degrading single-source detectors in unseen environments. Existing single-domain general…

cs.CV2026

Online Reasoning Video Object Segmentation

Jinyuan Liu, Yang Wang, Zeyu Zhao +3

Reasoning video object segmentation predicts pixel-level masks in videos from natural-language queries that may involve implicit and temporally grounded references. However, existi…

cs.CV2024

From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera Calibration

Zekun Qian, Ruize Han, Wei Feng +2

We tackle a new problem of multi-view camera and subject registration in the bird's eye view (BEV) without pre-given camera calibration. This is a very challenging problem since it…

cs.CV2026

Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference

Tong Jin, Yunpeng Liu, Shuyu Hu +4

Recent visual place recognition (VPR) methods based on vision transformers, particularly foundation models, have achieved remarkable recognition performance. However, these models…

cs.CV2026

CD-RMOT-Bench: Benchmarking the Cross-Domain Referring Multi-Object Tracking

Xiangqun Zhang, Likai Wang, Zekun Qian +2

The paper introduces CD-RMOT-Bench, a benchmark for evaluating how well referring multi-object tracking models trained on one visual domain perform on different, unseen domains, an…

#referring multi-object tracking#cross-domain adaptation#language-guided tracking#domain shift
cs.CV2026

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

Yupeng Zhang, Ruize Han, Fangnan Zhou +2

Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of shifts often occur simultaneously…

cs.CV2022

From Indoor To Outdoor: Unsupervised Domain Adaptive Gait Recognition

Likai Wang, Ruize Han, Wei Feng +1

Gait recognition is an important AI task, which has been progressed rapidly with the development of deep learning. However, existing learning based gait recognition methods mainly…

cs.CV2025

CLIPVehicle: A Unified Framework for Vision-based Vehicle Search

Likai Wang, Ruize Han, Xiangqun Zhang +1

Vehicles, as one of the most common and significant objects in the real world, the researches on which using computer vision technologies have made remarkable progress, such as veh…

cs.CV2026

COVTrack++: Learning Open-Vocabulary Multi-Object Tracking from Continuous Videos via a Synergistic Paradigm

Zekun Qian, Wei Feng, Ruize Han +1

Multi-Object Tracking (MOT) has traditionally focused on a few specific categories, restricting its applicability to real-world scenarios involving diverse objects. Open-Vocabulary…

cs.CV2023

Combining the Silhouette and Skeleton Data for Gait Recognition

Likai Wang, Ruize Han, Wei Feng

Gait recognition, a long-distance biometric technology, has aroused intense interest recently. Currently, the two dominant gait recognition works are appearance-based and model-bas…

cs.CV2019

Multiple Human Association between Top and Horizontal Views by Matching Subjects' Spatial Distributions

Ruize Han, Yujun Zhang, Wei Feng +5

Video surveillance can be significantly enhanced by using both top-view data, e.g., those from drone-mounted cameras in the air, and horizontal-view data, e.g., those from wearable…

cs.CV2022

Single Object Tracking Research: A Survey

Ruize Han, Wei Feng, Qing Guo +1

Visual object tracking is an important task in computer vision, which has many real-world applications, e.g., video surveillance, visual navigation. Visual object tracking also has…

cs.CV2026

The Fourth Challenge on Image Super-Resolution (4) at NTIRE 2026: Benchmark Results and Method Overview

Zheng Chen, Kai Liu, Jingkai Wang +150

This paper presents the NTIRE 2026 image super-resolution (4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to r…

cs.CV2026

LV-OSD: Language-Vision-Complementary Open-Set Object Detection

Yupeng Zhang, Ruize Han, Wei Feng +2

Object detection is an important task in computer vision, which aims to detect the objects of interest. through the given category list or query images. In this work, we propose a…

cs.CV2026

BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning

Zekun Qian, Ruize Han, Wei Feng

Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack exp…

cs.CV2024

OCTrack: Benchmarking the Open-Corpus Multi-Object Tracking

Zekun Qian, Ruize Han, Wei Feng +3

We study a novel yet practical problem of open-corpus multi-object tracking (OCMOT), which extends the MOT into localizing, associating, and recognizing generic-category objects of…

cs.AI2024

Robust Collaborative Perception without External Localization and Clock Devices

Zixing Lei, Zhenyang Ni, Ruize Han +5

A consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information excha…

cs.CV2026

ExDet: Open-Domain Open-Vocabulary Detection with Cross-modal Extrapolation and Rectification

Yupeng Zhang, Yuzhong Feng, Ruize Han +3

Open-domain open-vocabulary detection (ODOVD) requires detectors to generalize to both novel categories and unseen domains, making it more challenging than open-vocabulary detectio…

cs.CV2024

Unveiling the Power of Self-supervision for Multi-view Multi-human Association and Tracking

Wei Feng, Feifan Wang, Ruize Han +2

Multi-view multi-human association and tracking (MvMHAT), is a new but important problem for multi-person scene video surveillance, aiming to track a group of people over time in e…

cs.NI2022

Access Trends of In-network Cache for Scientific Data

Ruize Han, Alex Sim, Kesheng Wu +6

Scientific collaborations are increasingly relying on large volumes of data for their work and many of them employ tiered systems to replicate the data to their worldwide user comm…

cs.CV2022

Self-supervised Social Relation Representation for Human Group Detection

Jiacheng Li, Ruize Han, Haomin Yan +3

Human group detection, which splits crowd of people into groups, is an important step for video-based human social activity analysis. The core of human group detection is the human…

cs.CV2024

OVT-B: A New Large-Scale Benchmark for Open-Vocabulary Multi-Object Tracking

Haiji Liang, Ruize Han

Open-vocabulary object perception has become an important topic in artificial intelligence, which aims to identify objects with novel classes that have not been seen during trainin…

cs.CV2026

NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection

Yupeng Zhang, Ruize Han, Zhiwei Chen +2

Despite the remarkable progress in open-vocabulary object detection (OVD), a significant gap remains between the training and testing phases. During training, the RPN and RoI heads…

cs.CV2025

Synthetic-To-Real Video Person Re-ID

Xiangqun Zhang, Wei Feng, Ruize Han +3

Person re-identification (Re-ID) is an important task and has significant applications for public security and information forensics, which has progressed rapidly with the developm…

cs.CV2022

Panoramic Human Activity Recognition

Ruize Han, Haomin Yan, Jiacheng Li +3

To obtain a more comprehensive activity understanding for a crowded scene, in this paper, we propose a new problem of panoramic human activity recognition (PAR), which aims to simu…

cs.CV2026

Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

Ruiqi Wu, Bingliang Jiao, Ruize Han +6

Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory…

cs.CV2022

A Benchmark of Video-Based Clothes-Changing Person Re-Identification

Likai Wang, Xiangqun Zhang, Ruize Han +4

Person re-identification (Re-ID) is a classical computer vision task and has achieved great progress so far. Recently, long-term Re-ID with clothes-changing has attracted increasin…

cs.CV2026

Timage: A Generative Text-in-Image Paradigm for Fine-Tuning Vision-Language Models

Yifeng Wu, Huimin Huang, Ruiluo Wu +5

Multimodal Large Language Models (MLLMs) often lose track of the right image regions during fine-grained spatial reasoning, because a textual query rarely carries any explicit geom…

cs.LG2026

IMWM: Intuition Models Complement World Models for Latent Planning

Baoqi Gao, Ruize Han, Miao Wang +1

Planning with a learned latent world model is a promising route to control from raw pixels, but a strong world model alone is not enough. We show this experimentally: even with a p…

cs.CV2026

RT-SDGOD: Real-Time Single-Domain Generalized Object Detection

Yupeng Zhang, Fangzhuo Gao, Ruize Han +2

In real-world deployment under strict real-time constraints, weather and imaging variations induce significant distribution shifts, severely degrading detectors. Single-Domain Gene…

cs.CV2026

The First Challenge on Remote Sensing Infrared Image Super-Resolution at NTIRE 2026: Benchmark Results and Method Overview

Kai Liu, Haoyang Yue, Zeli Lin +65

This paper presents the NTIRE 2026 Remote Sensing Infrared Image Super-Resolution (x4) Challenge, one of the associated challenges of NTIRE 2026. The challenge aims to recover high…

cs.CV2026

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

Yupeng Zhang, Ruize Han, Yuzhong Feng +3

Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-world category spaces continual…