1 paper
Zuojin Tang, Bin Hu, Chenyang Zhao +3
Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input si…