11 citations · 11 across the 4 of their papers we have counts for
4 papers
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier +29
In this work, we discuss building performant Multimodal Large Language Models (MLLMs). In particular, we study the importance of various architecture components and data choices. T…
Compress3D: a Compressed Latent Space for 3D Generation from a Single Image
Bowen Zhang, Tianyu Yang, Yu Li +2
3D generation has witnessed significant advancements, yet efficiently producing high-quality 3D assets from a single image remains challenging. In this paper, we present a triplane…
Reconstructing 3D Human Pose from RGB-D Data with Occlusions
Bowen Dang, Xi Zhao, Bowen Zhang +1
We propose a new method to reconstruct the 3D human body from RGB-D images with occlusions. The foremost challenge is the incompleteness of the RGB-D data due to occlusions between…
Investigating the Learning Behaviour of In-context Learning: A Comparison with Supervised Learning
Xindi Wang, Yufei Wang, Can Xu +6
Large language models (LLMs) have shown remarkable capacity for in-context learning (ICL), where learning a new task from just a few training examples is done without being explici…