papers

Publications (20)

cs.CV2024

Evidence-based Match-status-Aware Gait Recognition for Out-of-Gallery Gait Identification

Heming Du, Chen Liu, Ming Wang +3

Existing gait recognition methods typically identify individuals based on the similarity between probe and gallery samples. However, these methods often neglect the fact that the g…

cs.CV2025

Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty Minimization

Bingqing Zhang, Zhuo Cao, Heming Du +4

Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and…

cs.MM2024

Diverse Sign Language Translation

Xin Shen, Lei Shen, Shaozu Yuan +3

Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language tr…

cs.AI2024

Perceive, Reflect, and Plan: Designing LLM Agent for Goal-Directed City Navigation without Instructions

Qingbin Zeng, Qinglong Yang, Shunan Dong +4

This paper considers a scenario in city navigation: an AI agent is provided with language descriptions of the goal location with respect to some well-known landmarks; By only obser…

cs.CV2022

SEFormer: Structure Embedding Transformer for 3D Object Detection

Xiaoyu Feng, Heming Du, Yueqi Duan +2

Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a key challenge to 3D object detection on point cloud. Recently, Transfo…

cs.CV2025

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

Zhuo Cao, Heming Du, Bingqing Zhang +3

Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes…

eess.IV2025

CMamba: Learned Image Compression with State Space Models

Zhuojie Wu, Heming Du, Shuyun Wang +4

Learned Image Compression (LIC) has explored various architectures, such as Convolutional Neural Networks (CNNs) and transformers, in modeling image content distributions in order…

stat.ML2026

FreDN: Spectral Disentanglement for Time Series Forecasting via Learnable Frequency Decomposition

Zhongde An, Jinhong You, Jiyanglin Li +4

Time series forecasting is essential in a wide range of real world applications. Recently, frequency-domain methods have attracted increasing interest for their ability to capture…

cs.AI2023

Divide and Ensemble: Progressively Learning for the Unknown

Hu Zhang, Xin Shen, Heming Du +10

In the wheat nutrient deficiencies classification challenge, we present the DividE and EnseMble (DEEM) method for progressive test data predictions. We find that (1) test images ar…

cs.CV2024

Affective Behaviour Analysis via Integrating Multi-Modal Knowledge

Wei Zhang, Feng Qiu, Chen Liu +4

Affective Behavior Analysis aims to facilitate technology emotionally smart, creating a world where devices can understand and react to our emotions as humans do. To comprehensivel…

cs.CV2023

RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel Segmentation

MD Wahiduzzaman Khan, Hongwei Sheng, Hu Zhang +11

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina f…

cs.IR2026

Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts

Bingqing Zhang, Zhuo Cao, Heming Du +4

Modern video-text retrieval (VTR) models excel on in-distribution benchmarks but are highly vulnerable to real-world query shifts, where the distribution of query data deviates fro…

cs.CV2025

Dynamic Orchestration of Multi-Agent System for Real-World Multi-Image Agricultural VQA

Yan Ke, Xin Yu, Heming Du +2

Agricultural visual question answering is essential for providing farmers and researchers with accurate and timely knowledge. However, many existing approaches are predominantly de…

cs.CV2021

VTNet: Visual Transformer Network for Object Goal Navigation

Heming Du, Xin Yu, Liang Zheng

Object goal navigation aims to steer an agent towards a target object based on observations of the agent. It is of pivotal importance to design effective visual representations of…

cs.CV2026

ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss

Jiaying Ying, Heming Du, Kaihao Zhang +2

Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and human-computer interaction.…

cs.CV2024

MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset

Xin Shen, Heming Du, Hongwei Sheng +9

Isolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developi…

cs.CV2023

When 3D Bounding-Box Meets SAM: Point Cloud Instance Segmentation with Weak-and-Noisy Supervision

Qingtao Yu, Heming Du, Chen Liu +1

Learning from bounding-boxes annotations has shown great potential in weakly-supervised 3D point cloud instance segmentation. However, we observed that existing methods would suffe…

cs.CV2020

Learning Object Relation Graph and Tentative Policy for Visual Navigation

Heming Du, Xin Yu, Liang Zheng

Target-driven visual navigation aims at navigating an agent towards a given target based on the observation of the agent. In this task, it is critical to learn informative visual r…

cs.CV2024

FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding

Zhuo Cao, Bingqing Zhang, Heming Du +3

Text-guided Video Temporal Grounding (VTG) aims to localize relevant segments in untrimmed videos based on textual descriptions, encompassing two subtasks: Moment Retrieval (MR) an…

cs.CV2024

TokenBinder: Text-Video Retrieval with One-to-Many Alignment Paradigm

Bingqing Zhang, Zhuo Cao, Heming Du +4

Text-Video Retrieval (TVR) methods typically match query-candidate pairs by aligning text and video features in coarse-grained, fine-grained, or combined (coarse-to-fine) manners.…