1 paper
Ainaz Eftekhar, Kuo-Hao Zeng, Jiafei Duan +3
Embodied AI models often employ off the shelf vision backbones like CLIP to encode their visual observations. Although such general purpose representations encode rich syntactic an…