activity
20222024
most citedSAM3D: Segment Anything in 3D Scenes

33 citations · 55 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CV20241 cited

Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Zhenyu Wang, Enze Xie, Aoxue Li +3

Despite significant advancements in text-to-image models for generating high-quality images, these methods still struggle to ensure the controllability of text prompts over images…

cs.AI20243 cited

A Survey of Reasoning with Foundation Models

Jiankai Sun, Chuanyang Zheng, Enze Xie +31

Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It…

cs.CV20235 cited

OV-PARTS: Towards Open-Vocabulary Part Segmentation

Meng Wei, Xiaoyu Yue, Wenwei Zhang +3

Segmenting and recognizing diverse object parts is a crucial ability in applications spanning various computer vision and robotic tasks. While significant progress has been made in…

cs.CV20233 cited

Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection

Chenming Zhu, Wenwei Zhang, Tai Wang +2

Point cloud-based open-vocabulary 3D object detection aims to detect 3D categories that do not have ground-truth annotations in the training set. It is extremely challenging becaus…

cs.CV202333 cited

SAM3D: Segment Anything in 3D Scenes

Yunhan Yang, Xiaoyang Wu, Tong He +2

In this work, we propose SAM3D, a novel framework that is able to predict masks in 3D point clouds by leveraging the Segment-Anything Model (SAM) in RGB images without further trai…

cs.CV20232 cited

TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale

Ziyun Zeng, Yixiao Ge, Zhan Tong +3

The ultimate goal for foundation models is realizing task-agnostic, i.e., supporting out-of-the-box usage without task-specific fine-tuning. Although breakthroughs have been made i…