activity
20182022
most citedVision-Language Pre-Training with Triple Contrastive Learning

14 citations · 19 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2022

OSCARS: An Outlier-Sensitive Content-Based Radiography Retrieval System

Xiaoyuan Guo, Jiali Duan, Saptarshi Purkayastha +3

Improving the retrieval relevance on noisy datasets is an emerging need for the curation of a large-scale clean dataset in the medical domain. While existing methods can be applied…

cs.CV20225 cited

Multi-modal Alignment using Representation Codebook

Jiali Duan, Liqun Chen, Son Tran +4

Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusi…

cs.CV202214 cited

Vision-Language Pre-Training with Triple Contrastive Learning

Jinyu Yang, Jiali Duan, Son Tran +6

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…

cs.CV2020

SLADE: A Self-Training Framework For Distance Metric Learning

Jiali Duan, Yen-Liang Lin, Son Tran +2

Most existing distance metric learning approaches use fully labeled data to learn the sample similarities in an embedding space. We present a self-training framework, SLADE, to imp…

cs.RO2019

Robot Learning via Human Adversarial Games

Jiali Duan, Qian Wang, Lerrel Pinto +2

Much work in robotics has focused on "human-in-the-loop" learning techniques that improve the efficiency of the learning process. However, these algorithms have made the strong ass…

cs.CV2018

An Interpretable Generative Model for Handwritten Digit Image Synthesis

Yao Zhu, Saksham Suri, Pranav Kulkarni +3

An interpretable generative model for handwritten digits synthesis is proposed in this work. Modern image generative models, such as Generative Adversarial Networks (GANs) and Vari…