1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 1 cited
3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment
Ziyu Zhu, Xiaojian Ma, Yixin Chen +3
3D vision-language grounding (3D-VL) is an emerging field that aims to connect the 3D physical world with natural language, which is crucial for achieving embodied intelligence. Cu…
cs.CV2023
Improving Scene Graph Generation with Superpixel-Based Interaction Learning
Jingyi Wang, Can Zhang, Jinfa Huang +2
Recent advances in Scene Graph Generation (SGG) typically model the relationships among entities utilizing box-level features from pre-defined detectors. We argue that an overlooke…