125 citations · 179 across the 8 of their papers we have counts for
7 papers · 1 filter
ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Yucheng Han, Chi Zhang, Xin Chen +5
Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for spe…
Efficient Single-Image Depth Estimation on Mobile Devices, Mobile AI & AIM 2022 Challenge: Report
Andrey Ignatov, Grigory Malivenko, Radu Timofte +36
Various depth estimation models are now widely used on many mobile and IoT devices for image segmentation, bokeh effect rendering, object tracking and many other mobile tasks. Thus…
Learning Variational Motion Prior for Video-based Motion Capture
Xin Chen, Zhuo Su, Lingbo Yang +4
Motion capture from a monocular video is fundamental and crucial for us humans to naturally experience and interact with each other in Virtual Reality (VR) and Augmented Reality (A…
Hierarchical Normalization for Robust Monocular Depth Estimation
Chi Zhang, Wei Yin, Zhibin Wang +3
In this paper, we address monocular depth estimation with deep neural networks. To enable training of deep monocular estimation models with various sources of datasets, state-of-th…
Fine-grained Identity Preserving Landmark Synthesis for Face Reenactment
Haichao Zhang, Youcheng Ben, Weixi Zhang +3
Recent face reenactment works are limited by the coarse reference landmarks, leading to unsatisfactory identity preserving performance due to the distribution gap between the manip…
Shuffle Transformer with Feature Alignment for Video Face Parsing
Rui Zhang, Yang Han, Zilong Huang +4
This is a short technical report introducing the solution of the Team TCParser for Short-video Face Parsing Track of The 3rd Person in Context (PIC) Workshop and Challenge at CVPR…