Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
Hao Luo, Bohan Zhou, Zongqing Lu
Pre-training for Reinforcement Learning (RL) with purely video data is a valuable yet challenging problem. Although in-the-wild videos are readily available and inhere a vast amoun…
cs.CV2024
UniCode: Learning a Unified Codebook for Multimodal Large Language Models
Sipeng Zheng, Bohan Zhou, Yicheng Feng +2
In this paper, we propose \textbf{UniCode}, a novel approach within the domain of multimodal large language models (MLLMs) that learns a unified codebook to efficiently tokenize vi…