1 citations · 1 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Plan2Map: A Multimodal Benchmark for Document-Grounded Geospatial Boundary Reconstruction from Planning Records
Fabian Degen, Oishi Deb, Jindong Gu +4
Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machine-readable boundaries. We intr…
cs.CV2025
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
Yixuan Wu, Yang Zhang, Jian Wu +2
Multimodal Large Language Models (MLLMs) excel in vision-language tasks, such as image captioning and visual question answering. However, they often suffer from over-reliance on sp…
cs.CV2024★ 8 cited
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
Xianzheng Ma, Brandon Smart, Yash Bhalgat +14
As large language models (LLMs) evolve, their integration with 3D spatial data (3D-LLMs) has seen rapid progress, offering unprecedented capabilities for understanding and interact…