most citedMM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment

1 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

cs.RO2025

Universal Actions for Enhanced Embodied Foundation Models

Jinliang Zheng, Jianxiong Li, Dongxiu Liu +7

Training on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticea…

cs.RO2024

Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning

Jianxiong Li, Zhihao Wang, Jinliang Zheng +8

Multimodal task specification is essential for enhanced robotic performance, where \textit{Cross-modality Alignment} enables the robot to holistically understand complex task instr…

cs.CV20241 cited

MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment

Jihao Liu, Xin Huang, Jinliang Zheng +5

This paper introduces MM-Instruct, a large-scale dataset of diverse and high-quality visual instruction data designed to enhance the instruction-following capabilities of large mul…

cs.CV20241 cited

Enhancing Vision-Language Model with Unmasked Token Alignment

Jihao Liu, Jinliang Zheng, Boxiao Liu +2

Contrastive pre-training on image-text pairs, exemplified by CLIP, becomes a standard technique for learning multi-modal visual-language representations. Although CLIP has demonstr…

cs.CV2024

Instruction-Guided Visual Masking

Jinliang Zheng, Jianxiong Li, Sijie Cheng +6

Instruction following is crucial in contemporary LLM. However, when extended to multimodal setting, it often suffers from misalignment between specific textual instruction and targ…