activity
20202022
most citedVLUE: A Multi-Task Benchmark for Evaluating Vision-Language Models

4 citations · 4 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CL2022

EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning

Tiannan Wang, Wangchunshu Zhou, Yan Zeng +1

Pre-trained vision-language models (VLMs) have achieved impressive results in a range of vision-language tasks. However, popular VLMs usually consist of hundreds of millions of par…

cs.CV20224 cited

VLUE: A Multi-Task Benchmark for Evaluating Vision-Language Models

Wangchunshu Zhou, Yan Zeng, Shizhe Diao +1

Recent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) tasks. However, there exist several challenges for…

cond-mat.mtrl-sci2021

A probabilistic deep learning approach to automate the interpretation of multi-phase diffraction spectra

Nathan J. Szymanski, Christopher J. Bartel, Yan Zeng +2

Autonomous synthesis and characterization of inorganic materials requires the automatic and accurate analysis of X-ray diffraction spectra. For this task, we designed a probabilist…

cs.CL2020

Open-Domain Dialogue Generation Based on Pre-trained Language Models

Yan Zeng, Jian-Yun Nie

Pre-trained language models have been successfully used in response generation for open-domain dialogue. Four main frameworks have been proposed: (1) Transformer-ED using Transform…

cs.CL2020

Multi-Domain Dialogue State Tracking based on State Graph

Yan Zeng, Jian-Yun Nie

We investigate the problem of multi-domain Dialogue State Tracking (DST) with open vocabulary, which aims to extract the state from the dialogue. Existing approaches usually concat…

cs.CL2020

Jointly Optimizing State Operation Prediction and Value Generation for Dialogue State Tracking

Yan Zeng, Jian-Yun Nie

We investigate the problem of multi-domain Dialogue State Tracking (DST) with open vocabulary. Existing approaches exploit BERT encoder and copy-based RNN decoder, where the encode…