17 citations · 23 across the 3 of their papers we have counts for
4 papers
A Case Study on Combining ASR and Visual Features for Generating Instructional Video Captions
Jack Hessel, Bo Pang, Zhenhai Zhu +1
Instructional videos get high-traffic on video sharing platforms, and prior work suggests that providing time-stamped, subtask annotations (e.g., "heat the oil in the pan") improve…
Something's Brewing! Early Prediction of Controversy-causing Posts from Discussion Features
Jack Hessel, Lillian Lee
Controversial posts are those that split the preferences of a community, receiving both significant positive and significant negative feedback. Our inclusion of the word "community…
Unsupervised Discovery of Multimodal Links in Multi-image, Multi-sentence Documents
Jack Hessel, Lillian Lee, David Mimno
Images and text co-occur constantly on the web, but explicit links between images and sentences (or other intra-document textual units) are often not present. We present algorithms…
Cats and Captions vs. Creators and the Clock: Comparing Multimodal Content to Context in Predicting Relative Popularity
Jack Hessel, Lillian Lee, David Mimno
The content of today's social media is becoming more and more rich, increasingly mixing text, images, videos, and audio. It is an intriguing research question to model the interpla…