33 citations · 104 across the 10 of their papers we have counts for
11 papers · 1 filter
GMM: Delving into Gradient Aware and Model Perceive Depth Mining for Monocular 3D Detection
Weixin Mao, Jinrong Yang, Zheng Ge +5
Depth perception is a crucial component of monoc-ular 3D detection tasks that typically involve ill-posed problems. In light of the success of sample mining techniques in 2D object…
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction
Rui Yang, Lin Song, Yanwei Li +4
This paper aims to efficiently enable Large Language Models (LLMs) to use multimodal tools. Advanced proprietary LLMs, such as ChatGPT and GPT-4, have shown great potential for too…
BoxSnake: Polygonal Instance Segmentation with Box Supervision
Rui Yang, Lin Song, Yixiao Ge +1
Box-supervised instance segmentation has gained much attention as it requires only simple box annotations instead of costly mask or polygon annotations. However, existing box-super…
Dynamic Grained Encoder for Vision Transformers
Lin Song, Songyang Zhang, Songtao Liu +5
Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the…
Workshop on Autonomous Driving at CVPR 2021: Technical Report for Streaming Perception Challenge
Songyang Zhang, Lin Song, Songtao Liu +4
In this report, we introduce our real-time 2D object detection system for the realistic autonomous driving scenario. Our detector is built on a newly designed YOLO model, called YO…
Fine-Grained Dynamic Head for Object Detection
Lin Song, Yanwei Li, Zhengkai Jiang +4
The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, th…