62 citations · 265 across the 33 of their papers we have counts for
51 papers · 1 filter
IS-Fusion: Instance-Scene Collaborative Fusion for Multimodal 3D Object Detection
Junbo Yin, Jianbing Shen, Runnan Chen +4
Bird's eye view (BEV) representation has emerged as a dominant solution for describing 3D space in autonomous driving scenarios. However, objects in the BEV representation typicall…
ClusterFormer: Clustering As A Universal Visual Learner
James C. Liang, Yiming Cui, Qifan Wang +3
This paper presents CLUSTERFORMER, a universal vision model that is based on the CLUSTERing paradigm with TransFORMER. It comprises two novel designs: 1. recurrent cross-attention…
LOGICSEG: Parsing Visual Semantics with Neural Logic Learning and Reasoning
Liulei Li, Wenguan Wang, Yi Yang
Current high-performance semantic segmentation models are purely data-driven sub-symbolic approaches and blind to the structured nature of the visual world. This is in stark contra…
Logic-induced Diagnostic Reasoning for Semi-supervised Semantic Segmentation
Chen Liang, Wenguan Wang, Jiaxu Miao +1
Recent advances in semi-supervised semantic segmentation have been heavily reliant on pseudo labeling to compensate for limited labeled data, disregarding the valuable relational k…
Omnidirectional Information Gathering for Knowledge Transfer-based Audio-Visual Navigation
Jinyu Chen, Wenguan Wang, Si Liu +2
Audio-visual navigation is an audio-targeted wayfinding task where a robot agent is entailed to travel a never-before-seen 3D environment towards the sounding source. In this artic…
DREAMWALKER: Mental Planning for Continuous Vision-Language Navigation
Hanqing Wang, Wei Liang, Luc Van Gool +1
VLN-CE is a recently released embodied task, where AI agents need to navigate a freely traversable environment to reach a distant target location, given language instructions. It p…