2 papers
cs.CV2026
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Xiang An, Yin Xie, Feilong Tang +27
We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of mu…
cs.CV2026
DiffST: Spatiotemporal-Aware Diffusion for Real-World Space-Time Video Super-Resolution
Zheng Chen, Ruofan Yang, Jin Han +5
Diffusion-based models have shown strong performance in video super-resolution (VSR) and video frame interpolation (VFI). However, their role in the coupled space-time video super-…