activity
20212023
most citedTowards Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong Baseline

3 citations · 9 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2023★ 1 cited

Distortion-aware Transformer in 360° Salient Object Detection

Yinjie Zhao, Lichen Zhao, Qian Yu +3

With the emergence of VR and AR, 360° data attracts increasing attention from the computer vision and multimedia communities. Typically, 360° data is projected into 2D ERP (equirec…

cs.CV2023★ 3 cited

VL-SAT: Visual-Linguistic Semantics Assisted Training for 3D Semantic Scene Graph Prediction in Point Cloud

Ziqin Wang, Bowen Cheng, Lichen Zhao +3

The task of 3D semantic scene graph (3DSSG) prediction in the point cloud is challenging since (1) the 3D point cloud only captures geometric structures with limited semantics comp…

cs.CV2022★ 3 cited

Towards Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong Baseline

Lichen Zhao, Daigang Cai, Jing Zhang +6

Recently, 3D vision-and-language tasks have attracted increasing research interest. Compared to other vision-and-language tasks, the 3D visual question answering (VQA) task is less…

cs.CV2022★ 2 cited

Democratizing Contrastive Language-Image Pre-training: A CLIP Benchmark of Data, Model, and Supervision

Yufeng Cui, Lichen Zhao, Feng Liang +2

Contrastive Language-Image Pretraining (CLIP) has emerged as a novel paradigm to learn visual models from language supervision. While researchers continue to push the frontier of C…

cs.CV2021

Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Yangguang Li, Feng Liang, Lichen Zhao +5

Recently, large-scale Contrastive Language-Image Pre-training (CLIP) has attracted unprecedented attention for its impressive zero-shot recognition ability and excellent transferab…