74 citations · 169 across the 6 of their papers we have counts for
8 papers
SEED: Self-supervised Distillation For Visual Representation
Zhiyuan Fang, Jianfeng Wang, Lijuan Wang +3
This paper is concerned with self-supervised learning for small models. The problem is motivated by our empirical studies that while the widely used contrastive self-supervised lea…
Weak Supervision and Referring Attention for Temporal-Textual Association Learning
Zhiyuan Fang, Shu Kong, Zhe Wang +2
A system capturing the association between video frames and textual queries offer great potential for better video analysis. However, training such a system in a fully supervised w…
ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language
Zhe Wang, Zhiyuan Fang, Jun Wang +1
Person search by natural language aims at retrieving a specific person in a large-scale image pool that matches the given textual descriptions. While most of the current methods tr…
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee +2
Captioning is a crucial and challenging task for video understanding. In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the…
Modularized Textual Grounding for Counterfactual Resilience
Zhiyuan Fang, Shu Kong, Charless Fowlkes +1
Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding p…
Weakly Supervised Attention Learning for Textual Phrases Grounding
Zhiyuan Fang, Shu Kong, Tianshu Yu +1
Grounding textual phrases in visual content is a meaningful yet challenging problem with various potential applications such as image-text inference or text-driven multimedia inter…