most citedGrounded SAM: Assembling Open-World Models for Diverse Visual Tasks

93 citations · 133 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CV20243 cited

OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation

Bohao Peng, Xiaoyang Wu, Li Jiang +4

The booming of 3D recognition in the 2020s began with the introduction of point cloud transformers. They quickly overwhelmed sparse CNNs and became state-of-the-art models, especia…

cs.CL2024

E^2-LLM: Efficient and Extreme Length Extension of Large Language Models

Jiaheng Liu, Zhiqi Bai, Yuanxing Zhang +11

Typically, training LLMs with long context sizes is computationally expensive, requiring extensive training hours and GPU resources. Existing long-context extension methods usually…

cs.CV202493 cited

Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Tianhe Ren, Shilong Liu, Ailing Zeng +14

We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and seg…

cs.LG20236 cited

Data Pruning via Moving-one-Sample-out

Haoru Tan, Sitong Wu, Fei Du +4

In this paper, we propose a novel data-pruning approach called moving-one-sample-out (MoSo), which aims to identify and remove the least informative samples from the training set.…

cs.CV20234 cited

Mask-Attention-Free Transformer for 3D Instance Segmentation

Xin Lai, Yuhui Yuan, Ruihang Chu +3

Recently, transformer-based methods have dominated 3D instance segmentation, where mask attention is commonly involved. Specifically, object queries are guided by the initial insta…

cs.CV20235 cited

FocalFormer3D : Focusing on Hard Instance for 3D Object Detection

Yilun Chen, Zhiding Yu, Yukang Chen +4

False negatives (FN) in 3D object detection, {\em e.g.}, missing predictions of pedestrians, vehicles, or other obstacles, can lead to potentially dangerous situations in autonomou…