1 paper · 1 filter
Shangqing Tu, Yucheng Wang, Daniel Zhang-Li +8
Existing Large Vision-Language Models (LVLMs) can process inputs with context lengths up to 128k visual and text tokens, yet they struggle to generate coherent outputs beyond 1,000…