output
20172024
most citedRhythmNet: End-to-end Heart Rate Estimation from Face via Spatial-temporal Representation

421 citations

Showing cs.CVShow all

23 papers · 1 filter

cs.CV202448 cited

Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training

Man Yao, Xuerui Qiu, Tianxiang Hu +7

The ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major c…

cs.CV202415 cited

Contrastive Learning of Person-independent Representations for Facial Action Unit Detection

Yong Li, Shiguang Shan

Facial action unit (AU) detection, aiming to classify AU present in the facial image, has long suffered from insufficient AU annotations. In this paper, we aim to mitigate this dat…

cs.CV202312 cited

Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric Synthesis

Linsen Song, Wayne Wu, Chaoyou Fu +2

Existing automated dubbing methods are usually designed for Professionally Generated Content (PGC) production, which requires massive training data and training time to learn a per…

cs.CV202338 cited

End-to-end Alternating Optimization for Real-World Blind Super Resolution

Zhengxiong Luo, Yan Huang, Shang Li +2

Blind Super-Resolution (SR) usually involves two sub-problems: 1) estimating the degradation of the given low-resolution (LR) image; 2) super-resolving the LR image to its high-res…

cs.CV202240 cited

QueryProp: Object Query Propagation for High-Performance Video Object Detection

Fei He, Naiyu Gao, Jian Jia +2

Video object detection has been an important yet challenging topic in computer vision. Traditional methods mainly focus on designing the image-level or box-level feature propagatio…

cs.CV202297 cited

Weakly Aligned Feature Fusion for Multimodal Object Detection

Lu Zhang, Zhiyong Liu, Xiangyu Zhu +4

To achieve accurate and robust object detection in the real-world scenario, various forms of images are incorporated, such as color, thermal, and depth. However, multimodal data of…