activity
20182024
most citedWhen Does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?

6 citations · 11 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Lijie Fan, Tianhong Li, Siyang Qin +6

Scaling up autoregressive models in vision has not proven as beneficial as in large language models. In this work, we investigate this scaling problem in the context of text-to-ima…

cs.CV20216 cited

When Does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?

Lijie Fan, Sijia Liu, Pin-Yu Chen +2

Contrastive learning (CL) can learn generalizable feature representations and achieve the state-of-the-art performance of downstream tasks by finetuning a linear classifier on top…

cs.CV20205 cited

In-Home Daily-Life Captioning Using Radio Signals

Lijie Fan, Tianhong Li, Yuan Yuan +1

This paper aims to caption daily life --i.e., to create a textual description of people's activities and interactions with objects in their homes. Addressing this problem requires…

cs.CV2020

Learning Longterm Representations for Person Re-Identification Using Radio Signals

Lijie Fan, Tianhong Li, Rongyao Fang +3

Person Re-Identification (ReID) aims to recognize a person-of-interest across different places and times. Existing ReID methods rely on images or videos collected using RGB cameras…

cs.CV2019

Making the Invisible Visible: Action Recognition Through Walls and Occlusions

Tianhong Li, Lijie Fan, Mingmin Zhao +2

Understanding people's actions and interactions typically depends on seeing them. Automating the process of action recognition from visual data has been the topic of much research…

cs.CV2018

Controllable Image-to-Video Translation: A Case Study on Facial Expression Generation

Lijie Fan, Wenbing Huang, Chuang Gan +2

The recent advances in deep learning have made it possible to generate photo-realistic images by using neural networks and even to extrapolate video frames from an input video clip…