most citedWanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

8 citations · 20 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2023

Watermarking LLMs with Weight Quantization

Linyang Li, Botian Jiang, Pengyu Wang +3

Abuse of large language models reveals high risks as large language models are being deployed at an astonishing speed. It is important to protect the model weights to avoid malicio…

cs.CL20238 cited

WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Conghui He, Zhenjiang Jin, Chao Xu +6

The rise in popularity of ChatGPT and GPT-4 has significantly accelerated the development of large models, leading to the creation of numerous impressive large language models(LLMs…

cs.CL20235 cited

Does Correction Remain A Problem For Large Language Models?

Xiaowu Zhang, Xiaotian Zhang, Cheng Yang +2

As large language models, such as GPT, continue to advance the capabilities of natural language processing (NLP), the question arises: does the problem of correction still persist?…

cs.CL20234 cited

PromptNER: A Prompting Method for Few-shot Named Entity Recognition via k Nearest Neighbor Search

Mozhi Zhang, Hang Yan, Yaqian Zhou +1

Few-shot Named Entity Recognition (NER) is a task aiming to identify named entities via limited annotated samples. Recently, prototypical networks have shown promising performance…

cs.CL20233 cited

Unified Demonstration Retriever for In-Context Learning

Xiaonan Li, Kai Lv, Hang Yan +6

In-context learning is a new learning paradigm where a language model conditions on a few input-output pairs (demonstrations) and a test input, and directly outputs the prediction.…

cs.CV2016

Multi-way Particle Swarm Fusion

Chen Liu, Hang Yan, Pushmeet Kohli +1

This paper proposes a novel MAP inference framework for Markov Random Field (MRF) in parallel computing environments. The inference framework, dubbed Swarm Fusion, is a natural gen…