33 citations · 55 across the 11 of their papers we have counts for
11 papers
Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation
Zhenyu Wang, Enze Xie, Aoxue Li +3
Despite significant advancements in text-to-image models for generating high-quality images, these methods still struggle to ensure the controllability of text prompts over images…
A Survey of Reasoning with Foundation Models
Jiankai Sun, Chuanyang Zheng, Enze Xie +31
Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It…
OV-PARTS: Towards Open-Vocabulary Part Segmentation
Meng Wei, Xiaoyu Yue, Wenwei Zhang +3
Segmenting and recognizing diverse object parts is a crucial ability in applications spanning various computer vision and robotic tasks. While significant progress has been made in…
Object2Scene: Putting Objects in Context for Open-Vocabulary 3D Detection
Chenming Zhu, Wenwei Zhang, Tai Wang +2
Point cloud-based open-vocabulary 3D object detection aims to detect 3D categories that do not have ground-truth annotations in the training set. It is extremely challenging becaus…
SAM3D: Segment Anything in 3D Scenes
Yunhan Yang, Xiaoyang Wu, Tong He +2
In this work, we propose SAM3D, a novel framework that is able to predict masks in 3D point clouds by leveraging the Segment-Anything Model (SAM) in RGB images without further trai…
TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale
Ziyun Zeng, Yixiao Ge, Zhan Tong +3
The ultimate goal for foundation models is realizing task-agnostic, i.e., supporting out-of-the-box usage without task-specific fine-tuning. Although breakthroughs have been made i…