papers

Publications (8)

cs.CV2022

PromptDet: Towards Open-vocabulary Detection using Uncurated Images

Chengjian Feng, Yujie Zhong, Zequn Jie +5

The goal of this work is to establish a scalable pipeline for expanding an object detector towards novel/unseen categories, using zero manual annotations. To achieve that, we make…

cs.CV2022

Two-Stream Networks for Object Segmentation in Videos

Hannan Lu, Zhi Tian, Lirong Yang +2

Existing matching-based approaches perform video object segmentation (VOS) via retrieving support features from a pixel-level memory, while some pixels may suffer from lack of corr…

cs.CV2021

SwiftNet: Real-time Video Object Segmentation

Haochen Wang, Xiaolong Jiang, Haibing Ren +2

In this work we present SwiftNet for real-time semisupervised video object segmentation (one-shot VOS), which reports 77.8% J &F and 70 FPS on DAVIS 2017 validation dataset, leadin…

cs.CV2022

Target-Driven Structured Transformer Planner for Vision-Language Navigation

Yusheng Zhao, Jinyu Chen, Chen Gao +5

Vision-language navigation is the task of directing an embodied agent to navigate in 3D scenes with natural language instructions. For the agent, inferring the long-term navigation…

cs.CV2022

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

Junyu Luo, Jiahui Fu, Xianghao Kong +5

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage par…

cs.RO2023

Active Neural Topological Mapping for Multi-Agent Exploration

Xinyi Yang, Yuxiang Yang, Chao Yu +5

This paper investigates the multi-agent cooperative exploration problem, which requires multiple agents to explore an unseen environment via sensory signals in a limited time. A po…