1 paper
Chunhao Lu, Qiang Lu, Meichen Dong +1
Current end-to-end multi-modal models utilize different encoders and decoders to process input and output information. This separation hinders the joint representation learning of…