57 citations · 90 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 25 cited
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Yucheng Han, Chi Zhang, Xin Chen +5
Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for spe…
cs.CV2022★ 8 cited
Hierarchical Normalization for Robust Monocular Depth Estimation
Chi Zhang, Wei Yin, Zhibin Wang +3
In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-th…
cs.CV2021★ 57 cited
TFPose: Direct Human Pose Estimation with Transformers
Weian Mao, Yongtao Ge, Chunhua Shen +3
We propose a human pose estimation framework that solves the task in the regression-based fashion. Unlike previous regression-based methods, which often fall behind those state-of-…