4 papers
Human Pose Regression with Residual Log-likelihood Estimation
Jiefeng Li, Siyuan Bian, Ailing Zeng +4
Heatmap-based methods dominate in the field of human pose estimation by modelling the output distribution through likelihood heatmaps. In contrast, regression-based methods are mor…
Beyond Instructional Videos: Probing for More Diverse Visual-Textual Grounding on YouTube
Jack Hessel, Zhenhai Zhu, Bo Pang +1
Pretraining from unlabelled web videos has quickly become the de-facto means of achieving high performance on many video understanding tasks. Features are learned via prediction of…
A Case Study on Combining ASR and Visual Features for Generating Instructional Video Captions
Jack Hessel, Bo Pang, Zhenhai Zhu +1
Instructional videos get high-traffic on video sharing platforms, and prior work suggests that providing time-stamped, subtask annotations (e.g., "heat the oil in the pan") improve…
Decoupled Box Proposal and Featurization with Ultrafine-Grained Semantic Labels Improve Image Captioning and Visual Question Answering
Soravit Changpinyo, Bo Pang, Piyush Sharma +1
Object detection plays an important role in current solutions to vision and language tasks like image captioning and visual question answering. However, popular models like Faster…