activity
20192024
most citedImproving Multi-turn Dialogue Modelling with Utterance ReWriter

17 citations · 55 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CL2024

ProgGen: Generating Named Entity Recognition Datasets Step-by-step with Self-Reflexive Large Language Models

Yuzhao Heng, Chunyuan Deng, Yitong Li +4

Although Large Language Models (LLMs) exhibit remarkable adaptability across domains, these models often fall short in structured knowledge extraction tasks such as named entity re…

cs.CL2024★ 1 cited

TPD: Enhancing Student Language Model Reasoning via Principle Discovery and Guidance

Haorui Wang, Rongzhi Zhang, Yinghao Li +4

Large Language Models (LLMs) have recently showcased remarkable reasoning abilities. However, larger models often surpass their smaller counterparts in reasoning tasks, posing the…

cs.LG2023★ 4 cited

Local Boosting for Weakly-Supervised Learning

Rongzhi Zhang, Yue Yu, Jiaming Shen +2

Boosting is a commonly used technique to enhance the performance of a set of base models by combining them into a strong ensemble model. Though widely adopted, boosting is typicall…

cs.CL2023★ 2 cited

ReGen: Zero-Shot Text Classification via Training Data Generation with Progressive Dense Retrieval

Yue Yu, Yuchen Zhuang, Rongzhi Zhang +3

With the development of large language models (LLMs), zero-shot learning has attracted much attention for various NLP tasks. Different from prior works that generate training data…

cs.LG2023★ 5 cited

Do Not Blindly Imitate the Teacher: Using Perturbed Loss for Knowledge Distillation

Rongzhi Zhang, Jiaming Shen, Tianqi Liu +4

Knowledge distillation is a popular technique to transfer knowledge from large teacher models to a small student model. Typically, the student learns to imitate the teacher by mini…

cs.CL2022★ 7 cited

Cold-Start Data Selection for Few-shot Language Model Fine-tuning: A Prompt-Based Uncertainty Propagation Approach

Yue Yu, Rongzhi Zhang, Ran Xu +3

Large Language Models have demonstrated remarkable few-shot performance, but the performance can be sensitive to the selection of few-shot instances. We propose PATRON, a new metho…