2 papers
cs.CV2026
Twins: Learn to Predict Unified Representations with Focal Loss
Kaixiong Gong, Xin Cai, Bin Lin +9
Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the interface via a shared codeb…
cs.CV2026
STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching
Hao Wang, Haoran Geng, Xiaotong Yang +7
Stereo matching is a fundamental task in 3D reconstruction. Despite remarkable advances, the prevailing paradigms formulate stereo matching as a deterministic regression problem, c…