6 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 1 cited
Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
Zuyan Liu, Yuhao Dong, Yongming Rao +2
In the realm of vision-language understanding, the proficiency of models in interpreting and reasoning over visual content has become a cornerstone for numerous applications. Howev…
cs.CV2023
HandMIM: Pose-Aware Self-Supervised Learning for 3D Hand Mesh Estimation
Zuyan Liu, Gaojie Lin, Congyi Wang +2
With an enormous number of hand images generated over time, unleashing pose knowledge from unlabeled images for supervised hand mesh estimation is an emerging yet challenging topic…
cs.CV2023★ 6 cited
Unleashing Text-to-Image Diffusion Models for Visual Perception
Wenliang Zhao, Yongming Rao, Zuyan Liu +3
Diffusion models (DMs) have become the new trend of generative models and have demonstrated a powerful ability of conditional synthesis. Among those, text-to-image diffusion models…