1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
ROOT: VLM based System for Indoor Scene Understanding and Beyond
Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou +4
Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In…
cs.CV2024
TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
Bozhi Luan, Hao Feng, Hong Chen +3
The advent of Large Multimodal Models (LMMs) has sparked a surge in research aimed at harnessing their remarkable reasoning abilities. However, for understanding text-rich images,…
cs.CV2023
Progressive Recurrent Network for Shadow Removal
Yonghui Wang, Wengang Zhou, Hao Feng +2
Single-image shadow removal is a significant task that is still unresolved. Most existing deep learning-based approaches attempt to remove the shadow directly, which can not deal w…