3d reconstruction 1gaussian splatting 1multimodal generation 1panoramic video 1spatial world modeling 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
Mingchao Sun, Luyang Tang, Yu Liu +34
The paper introduces ABot-3DWorld 0, a multimodal system that converts text, images, or video into high‑fidelity, explorable 3D worlds using a compact spatial representation and pa…
cs.CV2026
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning
Chengzhen Yu, Canran Xiao, Siyuan Ma +1
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly con…