1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Yukang Chen, Fuzhao Xue, Dacheng Li +15
Long-context capability is critical for multi-modal foundation models, especially for long video understanding. We introduce LongVILA, a full-stack solution for long-context visual…