Publications (8)
PromptDet: Towards Open-vocabulary Detection using Uncurated Images
Chengjian Feng, Yujie Zhong, Zequn Jie +5
The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make…
Two-Stream Networks for Object Segmentation in Videos
Hannan Lu, Zhi Tian, Lirong Yang +2
Existing matching-based approaches perform video object segmentation (VOS) via retrieving support features from a pixel-level memory, while some pixels may suffer from lack of corr…
SwiftNet: Real-time Video Object Segmentation
Haochen Wang, Xiaolong Jiang, Haibing Ren +2
In this work we present SwiftNet for real-time semisupervised video object segmentation (one-shot VOS), which reports 77.8% J &F and 70 FPS on DAVIS 2017 validation dataset, leadin…
Target-Driven Structured Transformer Planner for Vision-Language Navigation
Yusheng Zhao, Jinyu Chen, Chen Gao +5
Vision-language navigation is the task of directing an embodied agent to navigate in 3D scenes with natural language instructions. For the agent, inferring the long-term navigation…
3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection
Junyu Luo, Jiahui Fu, Xianghao Kong +5
3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage par…
Active Neural Topological Mapping for Multi-Agent Exploration
Xinyi Yang, Yuxiang Yang, Chao Yu +5
This paper investigates the multi-agent cooperative exploration problem, which requires multiple agents to explore an unseen environment via sensory signals in a limited time. A po…