1 paper
Kaihang Pan, Siliang Tang, Juncheng Li +6
For multimodal LLMs, the synergy of visual comprehension (textual output) and generation (visual output) presents an ongoing challenge. This is due to a conflicting objective: for…