130 citations · 136 across the 6 of their papers we have counts for
9 papers · 1 filter
VideoGLUE: Video General Understanding Evaluation of Foundation Models
Liangzhe Yuan, Nitesh Bharadwaj Gundavarapu, Long Zhao +14
We evaluate the video understanding capabilities of existing foundation models (FMs) using a carefully designed experiment protocol consisting of three hallmark tasks (action recog…
Emergent Correspondence from Image Diffusion
Luming Tang, Menglin Jia, Qianqian Wang +2
Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explici…
Exploring Visual Engagement Signals for Representation Learning
Menglin Jia, Zuxuan Wu, Austin Reiter +3
Visual engagement in social media platforms comprises interactions with photo posts including comments, shares, and likes. In this paper, we leverage such visual engagement clues a…
Intentonomy: a Dataset and Study towards Human Intent Understanding
Menglin Jia, Zuxuan Wu, Austin Reiter +3
An image is worth a thousand words, conveying information that goes beyond the physical visual content therein. In this paper, we study the intent behind social media images with a…
Learning Occupancy Function from Point Clouds for Surface Reconstruction
Meng Jia, Matthew Kyan
Implicit function based surface reconstruction has been studied for a long time to recover 3D shapes from point clouds sampled from surfaces. Recently, Signed Distance Functions (S…
Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
Menglin Jia, Mengyun Shi, Mikhail Sirotenko +5
In this work we explore the task of instance segmentation with attribute localization, which unifies instance segmentation (detect and segment each object instance) and fine-graine…