5 papers
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
Chang Liu, Henghui Ding, Nikhila Ravi +40
This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, whi…
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
Jihwan Hong, Jaeyoung Do
Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe-based approaches that couple a v…
3rd Place of MeViS-Audio Track of the 5th PVUW: VIRST-Audio
Jihwan Hong, Jaeyoung Do
Audio-based Referring Video Object Segmentation (ARVOS) requires grounding audio queries into pixel-level object masks over time, posing challenges in bridging acoustic signals wit…
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo +5
Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image to…
Dynin-Omni: Omnimodal Unified Large Diffusion Language Model
Jaeik Kim, Woojin Kim, Jihwan Hong +8
We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understand…