2 papers
cs.CV2026
UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis
Sihan Chen, Xiang Zhang, Yang Zhang +2
With the recent surge of generative models, diffusion-based approaches have become mainstream for view synthesis tasks, either in an explicit depth-warp-inpaint or in an implicit e…
cs.LG2025
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset
Jing Liu, Sihan Chen, Xingjian He +4
In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-langu…