16 citations · 19 across the 20 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models
Shi-Yu Tian, Zhi Zhou, Kun-Yang Yu +5
Spatial reasoning is a cornerstone capability for intelligent systems to perceive and interact with the physical world. However, multimodal large language models (MLLMs) frequently…
cs.CV2024
You Only Submit One Image to Find the Most Suitable Generative Model
Zhi Zhou, Lan-Zhe Guo, Peng-Xiao Song +1
Deep generative models have achieved promising results in image generation, and various generative model hubs, e.g., Hugging Face and Civitai, have been developed that enable model…