1 paper · 1 filter
Yuan Sun, Zhao Zhang, Jorge Ortiz
In current multimodal tasks, models typically freeze the encoder and decoder while adapting intermediate layers to task-specific goals, such as region captioning. Region-level visu…