activity
20232025
most citedUni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

3 citations · 3 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CV2025

3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o

Dingning Liu, Cheng Wang, Peng Gao +4

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existi…

cs.CV2024

TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction

Xuying Zhang, Yutong Liu, Yangguang Li +8

We present TAR3D, a novel framework that consists of a 3D-aware Vector Quantized-Variational AutoEncoder (VQ-VAE) and a Generative Pre-trained Transformer (GPT) to generate high-qu…

cs.CV20243 cited

Uni3D-LLM: Unifying Point Cloud Perception, Generation and Editing with Large Language Models

Dingning Liu, Xiaoshui Huang, Yuenan Hou +5

In this paper, we introduce Uni3D-LLM, a unified framework that leverages a Large Language Model (LLM) to integrate tasks of 3D perception, generation, and editing within point clo…

cs.CV2023

A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise

Chaoyou Fu, Renrui Zhang, Zihan Wang +15

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow L…

cs.AI2023

3DAxiesPrompts: Unleashing the 3D Spatial Task Capabilities of GPT-4V

Dingning Liu, Xiaomeng Dong, Renrui Zhang +5

In this work, we present a new visual prompting method called 3DAxiesPrompts (3DAP) to unleash the capabilities of GPT-4V in performing 3D spatial tasks. Our investigation reveals…