1 citations · 4 across the 11 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
Yiting Lu, Wei Luo, Peiyan Tu +8
World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct rea…
cs.CV2025
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
Bob Zhang, Haoran Li, Tao Zhang +5
Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instruct…
cs.CV2024★ 1 cited
ROOT: VLM based System for Indoor Scene Understanding and Beyond
Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou +4
Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In…