papers

Publications (30)

eess.IV2019

Encoding CT Anatomy Knowledge for Unpaired Chest X-ray Image Decomposition

Zeju Li, Han Li, Hu Han +3

Although chest X-ray (CXR) offers a 2D projection with overlapped anatomies, it is widely used for clinical diagnosis. There is clinical evidence supporting that decomposing an X-r…

cs.CV2023

Mixed-order self-paced curriculum learning for universal lesion detection

Han Li, Hu Han, S. Kevin Zhou

Self-paced curriculum learning (SCL) has demonstrated its great potential in computer vision, natural language processing, etc. During training, it implements easy-to-hard sampling…

cs.CV2022

Affective Behaviour Analysis Using Pretrained Model with Facial Priori

Yifan Li, Haomiao Sun, Zhaori Liu +1

Affective behaviour analysis has aroused researchers' attention due to its broad applications. However, it is labor exhaustive to obtain accurate annotations for massive face image…

cs.CV2021

Conditional Training with Bounding Map for Universal Lesion Detection

Han Li, Long Chen, Hu Han +1

Universal Lesion Detection (ULD) in computed tomography plays an essential role in computer-aided diagnosis. Promising ULD results have been reported by coarse-to-fine two-stage de…

cs.CV2017

Heterogeneous Face Attribute Estimation: A Deep Multi-Task Learning Approach

Hu Han, Anil K. Jain, Fang Wang +2

Face attribute estimation has many potential applications in video surveillance, face retrieval, and social media. While a number of methods have been proposed for face attribute e…

cs.AI2026

Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?

Kaiwen Zheng, Junchen Fu, Wenhao Deng +3

The paper introduces Light-MER, a sub‑billion‑parameter multimodal emotion recognition model that uses knowledge distillation, an optimal transport loss, and a multi‑reward optimiz…

#multimodal emotion recognition#knowledge distillation#lightweight models#optimal transport loss
cs.CV2024

Task-adaptive Q-Face

Haomiao Sun, Mingjie He, Shiguang Shan +2

Although face analysis has achieved remarkable improvements in the past few years, designing a multi-task face analysis model is still challenging. Most face analysis tasks are stu…

cond-mat.mes-hall2018

Enhancement of the thermoelectric effect due to the Majorana zero modes coupled to one quantum-dot system

Xiao-Qi Wang, Shu-Feng Zhang, Hu Han +2

By considering Majorana zero modes to laterally couple to the quantum dot, we evaluate the thermoelectric effect in one single-dot system. The calculation results show that if one…

cs.CV2023

Decoupled Textual Embeddings for Customized Image Generation

Yufei Cai, Yuxiang Wei, Zhilong Ji +3

Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suff…

cs.CV2021

Deep Learning to Segment Pelvic Bones: Large-scale CT Datasets and Baseline Models

Pengbo Liu, Hu Han, Yuanqi Du +9

Purpose: Pelvic bone segmentation in CT has always been an essential step in clinical diagnosis and surgery planning of pelvic bone diseases. Existing methods for pelvic bone segme…

cs.CV2020

Human Recognition Using Face in Computed Tomography

Jiuwen Zhu, Hu Han, S. Kevin Zhou

With the mushrooming use of computed tomography (CT) images in clinical decision making, management of CT data becomes increasingly difficult. From the patient identification persp…

cs.CV2020

Miss the Point: Targeted Adversarial Attack on Multiple Landmark Detection

Qingsong Yao, Zecheng He, Hu Han +1

Recent methods in multiple landmark detection based on deep convolutional neural networks (CNNs) reach high accuracy and improve traditional clinical workflow. However, the vulnera…

cs.CV2020

Cross-domain Face Presentation Attack Detection via Multi-domain Disentangled Representation Learning

Guoqing Wang, Hu Han, Shiguang Shan +1

Face presentation attack detection (PAD) has been an urgent problem to be solved in the face recognition systems. Conventional approaches usually assume the testing and training ar…

cs.CV2019

RhythmNet: End-to-end Heart Rate Estimation from Face via Spatial-temporal Representation

Xuesong Niu, Shiguang Shan, Hu Han +1

Heart rate (HR) is an important physiological signal that reflects the physical and emotional status of a person. Traditional HR measurements usually rely on contact monitors, whic…

cs.CV2020

Multi-label Co-regularization for Semi-supervised Facial Action Unit Recognition

Xuesong Niu, Hu Han, Shiguang Shan +1

Facial action units (AUs) recognition is essential for emotion analysis and has been widely applied in mental state analysis. Existing work on AU recognition usually requires big f…

cs.CV2019

FCSR-GAN: Joint Face Completion and Super-resolution via Multi-task Learning

Jiancheng Cai, Hu Han, Shiguang Shan +1

Combined variations containing low-resolution and occlusion often present in face images in the wild, e.g., under the scenario of video surveillance. While most of the existing fac…

cs.CV2022

Automatic Facial Paralysis Estimation with Facial Action Units

Xuri Ge, Joemon M. Jose, Pengcheng Wang +3

Facial palsy is unilateral facial nerve weakness or paralysis of rapid onset with unknown causes. Automatically estimating facial palsy severeness can be helpful for the diagnosis…

cs.CV2020

Bounding Maps for Universal Lesion Detection

Han Li, Hu Han, S. Kevin Zhou

Universal Lesion Detection (ULD) in computed tomography plays an essential role in computer-aided diagnosis systems. Many detection approaches achieve excellent results for ULD usi…

eess.IV2022

SATr: Slice Attention with Transformer for Universal Lesion Detection

Han Li, Long Chen, Hu Han +1

Universal Lesion Detection (ULD) in computed tomography plays an essential role in computer-aided diagnosis. Promising ULD results have been reported by multi-slice-input detection…

eess.IV2019

3D U-Net: A 3D Universal U-Net for Multi-Domain Medical Image Segmentation

Chao Huang, Hu Han, Qingsong Yao +2

Fully convolutional neural networks like U-Net have been the state-of-the-art methods in medical image segmentation. Practically, a network is highly specialized and trained separa…

cs.CV2020

Video-based Remote Physiological Measurement via Cross-verified Feature Disentangling

Xuesong Niu, Zitong Yu, Hu Han +3

Remote physiological measurements, e.g., remote photoplethysmography (rPPG) based heart rate (HR), heart rate variability (HRV) and respiration frequency (RF) measuring, are playin…

cs.CV2025

MoMBS: Mixed-order minibatch sampling enhances model training from diverse-quality images

Han Li, Hu Han, S. Kevin Zhou

Natural images exhibit label diversity (clean vs. noisy) in noisy-labeled image classification and prevalence diversity (abundant vs. sparse) in long-tailed image classification. S…

cs.CV2024

MGRR-Net: Multi-level Graph Relational Reasoning Network for Facial Action Units Detection

Xuri Ge, Joemon M. Jose, Songpei Xu +2

The Facial Action Coding System (FACS) encodes the action units (AUs) in facial images, which has attracted extensive research attention due to its wide use in facial expression an…

cs.CV2024

Face-MLLM: A Large Face Perception Model

Haomiao Sun, Mingjie He, Tianheng Lian +2

Although multimodal large language models (MLLMs) have achieved promising results on a wide range of vision-language tasks, their ability to perceive and understand human faces is…

cs.CV2018

VIPL-HR: A Multi-modal Database for Pulse Estimation from Less-constrained Face Video

Xuesong Niu, Hu Han, Shiguang Shan +1

Heart rate (HR) is an important physiological signal that reflects the physical and emotional activities of humans. Traditional HR measurements are mainly based on contact monitors…

cs.CV2021

Where is the disease? Semi-supervised pseudo-normality synthesis from an abnormal image

Yuanqi Du, Quan Quan, Hu Han +1

Pseudo-normality synthesis, which computationally generates a pseudo-normal image from an abnormal one (e.g., with lesions), is critical in many perspectives, from lesion detection…

cs.CV2018

Tattoo Image Search at Scale: Joint Detection and Compact Representation Learning

Hu Han, Jie Li, Anil K. Jain +2

The explosive growth of digital images in video surveillance and social media has led to the significant need for efficient search of persons of interest in law enforcement and for…

cs.CV2020

The 1st Challenge on Remote Physiological Signal Sensing (RePSS)

Xiaobai Li, Hu Han, Hao Lu +5

Remote measurement of physiological signals from videos is an emerging topic. The topic draws great interests, but the lack of publicly available benchmark databases and a fair val…

cs.CV2025

Dynamically evolving segment anything model with continuous learning for medical image segmentation

Zhaori Liu, Mengyang Li, Hu Han +3

Medical image segmentation is essential for clinical diagnosis, surgical planning, and treatment monitoring. Traditional approaches typically strive to tackle all medical image seg…

cs.CV2025

EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models

Yufei Cai, Hu Han, Yuxiang Wei +2

The progress on generative models has led to significant advances on text-to-video (T2V) generation, yet the motion controllability of generated videos remains limited. Existing mo…