12 citations · 27 across the 5 of their papers we have counts for
5 papers
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Yanwei Li, Yuechen Zhang, Chengyao Wang +5
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic…
DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
Tianqi Wang, Enze Xie, Ruihang Chu +2
End-to-end driving has made significant progress in recent years, demonstrating benefits such as system simplicity and competitive driving performance under both open-loop and clos…
A Survey of Reasoning with Foundation Models
Jiankai Sun, Chuanyang Zheng, Enze Xie +31
Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It…
Mask-Attention-Free Transformer for 3D Instance Segmentation
Xin Lai, Yuhui Yuan, Ruihang Chu +3
Recently, transformer-based methods have dominated 3D instance segmentation, where mask attention is commonly involved. Specifically, object queries are guided by the initial insta…
TriVol: Point Cloud Rendering via Triple Volumes
Tao Hu, Xiaogang Xu, Ruihang Chu +1
Existing learning-based methods for point cloud rendering adopt various 3D representations and feature querying mechanisms to alleviate the sparsity problem of point clouds. Howeve…