most citedDiffCap: Exploring Continuous Diffusion on Image Captioning

3 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2024

Improving Event Definition Following For Zero-Shot Event Detection

Zefan Cai, Po-Nien Kung, Ashima Suvarna +6

Existing approaches on zero-shot event detection usually train models on datasets annotated with known event types, and prompt them with unseen event definitions. These approaches…

cs.CL2024

PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Liang Chen, Yichi Zhang, Shuhuai Ren +7

We present PCA-Bench, a multimodal decision-making benchmark for evaluating the integrated capabilities of Multimodal Large Language Models (MLLMs). Departing from previous benchma…

cs.CV2024

VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness

Rongyu Zhang, Zefan Cai, Huanrui Yang +9

Finetuning a pretrained vision model (PVM) is a common technique for learning downstream vision tasks. However, the conventional finetuning process with randomly sampled data point…

cs.CL2023

DialogVCS: Robust Natural Language Understanding in Dialogue System Upgrade

Zefan Cai, Xin Zheng, Tianyu Liu +7

In the constant updates of the product dialogue systems, we need to retrain the natural language understanding (NLU) model as new data from the real users would be merged into the…

cs.CV20233 cited

DiffCap: Exploring Continuous Diffusion on Image Captioning

Yufeng He, Zefan Cai, Xu Gan +1

Current image captioning works usually focus on generating descriptions in an autoregressive manner. However, there are limited works that focus on generating descriptions non-auto…