Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Junxian Li, Kai Liu, Leyang Chen +7
Unified multimodal models (UMMs) have shown impressive capabilities in generating natural images and supporting multimodal reasoning. However, their potential in supporting compute…
cs.CV2025
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
Jiajie Guo, Qingpeng Zhu, Jin Zeng +3
Multimodal large language models (MLLMs) have achieved significant progress in image and language tasks due to the strong reasoning capability of large language models (LLMs). Neve…
cs.CV2024
Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription
Hongxiang Zhao, Xili Dai, Jianan Wang +5
Large image diffusion models have demonstrated zero-shot capability in novel view synthesis (NVS). However, existing diffusion-based NVS methods struggle to generate novel views th…