most citedMini-Gemini: Mining the Potential of Multi-modality Vision Language Models

12 citations · 27 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV202412 cited

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Yanwei Li, Yuechen Zhang, Chengyao Wang +5

In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic…

cs.CV20244 cited

DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving

Tianqi Wang, Enze Xie, Ruihang Chu +2

End-to-end driving has made significant progress in recent years, demonstrating benefits such as system simplicity and competitive driving performance under both open-loop and clos…

cs.AI20243 cited

A Survey of Reasoning with Foundation Models

Jiankai Sun, Chuanyang Zheng, Enze Xie +31

Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It…

cs.CV20234 cited

Mask-Attention-Free Transformer for 3D Instance Segmentation

Xin Lai, Yuhui Yuan, Ruihang Chu +3

Recently, transformer-based methods have dominated 3D instance segmentation, where mask attention is commonly involved. Specifically, object queries are guided by the initial insta…

cs.CV20234 cited

TriVol: Point Cloud Rendering via Triple Volumes

Tao Hu, Xiaogang Xu, Ruihang Chu +1

Existing learning-based methods for point cloud rendering adopt various 3D representations and feature querying mechanisms to alleviate the sparsity problem of point clouds. Howeve…