9 citations · 11 across the 5 of their papers we have counts for
5 papers · 1 filter
Align before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
Yifei Chen, Dapeng Chen, Ruijin Liu +3
Large-scale visual-language pre-trained models have achieved significant success in various video tasks. However, most existing methods follow an "adapt then align" paradigm, which…
PBFormer: Capturing Complex Scene Text Shape with Polynomial Band Transformer
Ruijin Liu, Ning Lu, Dapeng Chen +3
We present PBFormer, an efficient yet powerful scene text detector that unifies the transformer with a novel text shape representation Polynomial Band (PB). The representation has…
Video Action Recognition with Attentive Semantic Units
Yifei Chen, Dapeng Chen, Ruijin Liu +2
Visual-Language Models (VLMs) have significantly advanced action video recognition. Supervised by the semantics of action labels, recent works adapt the visual branch of VLMs to le…
Learning to Predict 3D Lane Shape and Camera Pose from a Single Image via Geometry Constraints
Ruijin Liu, Dapeng Chen, Tie Liu +2
Detecting 3D lanes from the camera is a rising problem for autonomous vehicles. In this task, the correct camera pose is the key to generating accurate lanes, which can transform a…
End-to-end Lane Shape Prediction with Transformers
Ruijin Liu, Zejian Yuan, Tie Liu +1
Lane detection, the process of identifying lane markings as approximated curves, is widely used for lane departure warning and adaptive cruise control in autonomous vehicles. The p…