120 citations · 373 across the 21 of their papers we have counts for
21 papers · 1 filter
GAFlow: Incorporating Gaussian Attention into Optical Flow
Ao Luo, Fan Yang, Xin Li +4
Optical flow, or the estimation of motion fields from image sequences, is one of the fundamental problems in computer vision. Unlike most pixel-wise tasks that aim at achieving con…
ContentCTR: Frame-level Live Streaming Click-Through Rate Prediction with Multimodal Transformer
Jiaxin Deng, Dong Shen, Shiyao Wang +4
In recent years, live streaming platforms have gained immense popularity as they allow users to broadcast their videos and interact in real-time with hosts and peers. Due to the dy…
Student Classroom Behavior Detection based on YOLOv7-BRA and Multi-Model Fusion
Fan Yang, Tao Wang, Xiaofei Wang
Accurately detecting student behavior in classroom videos can aid in analyzing their classroom performance and improving teaching effectiveness. However, the current accuracy rate…
SSD-MonoDETR: Supervised Scale-aware Deformable Transformer for Monocular 3D Object Detection
Xuan He, Fan Yang, Kailun Yang +5
Transformer-based methods have demonstrated superior performance for monocular 3D object detection recently, which aims at predicting 3D attributes from a single 2D image. Most exi…
Generation-Guided Multi-Level Unified Network for Video Grounding
Xing Cheng, Xiangyu Wu, Dong Shen +2
Video grounding aims to locate the timestamps best matching the query description within an untrimmed video. Prevalent methods can be divided into moment-level and clip-level frame…
Geo-Localization via Ground-to-Satellite Cross-View Image Retrieval
Zelong Zeng, Zheng Wang, Fan Yang +1
The large variation of viewpoint and irrelevant content around the target always hinder accurate image retrieval and its subsequent tasks. In this paper, we investigate an extremel…