1 paper
Jinbo Yan, Limeng Qiao, Jie Qin +3
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine…