4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 4 cited
Expression Prompt Collaboration Transformer for Universal Referring Video Object Segmentation
Jiajun Chen, Jiacheng Lin, Guojin Zhong +4
Audio-guided Video Object Segmentation (A-VOS) and Referring Video Object Segmentation (R-VOS) are two highly related tasks that both aim to segment specific objects from video seq…
cs.CV2023
SSD-MonoDETR: Supervised Scale-aware Deformable Transformer for Monocular 3D Object Detection
Xuan He, Fan Yang, Kailun Yang +5
Transformer-based methods have demonstrated superior performance for monocular 3D object detection recently, which aims at predicting 3D attributes from a single 2D image. Most exi…