4 papers
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
Haoyuan Li, Qihang Cao, Tao Tang +6
Recent progress in spatial reasoning with Multimodal Large Language Models (MLLMs) increasingly leverages geometric priors from 3D encoders. However, most existing integration stra…
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
Xiaojun Shan, Qi Cao, Xing Han +2
Recent advances in multimodal foundation models have achieved state-of-the-art performance across a range of tasks. These breakthroughs are largely driven by new pre-training parad…
Data Augmentation for Text-based Person Retrieval Using Large Language Models
Zheng Li, Lijia Si, Caili Guo +2
Text-based Person Retrieval (TPR) aims to retrieve person images that match the description given a text query. The performance improvement of the TPR model relies on high-quality…
Fine-Tuning InstructPix2Pix for Advanced Image Colorization
Zifeng An, Zijing Xu, Eric Fan +1
This paper presents a novel approach to human image colorization by fine-tuning the InstructPix2Pix model, which integrates a language model (GPT-3) with a text-to-image model (Sta…