activity
20212024
most citedAugGPT: Leveraging ChatGPT for Text Data Augmentation

99 citations · 240 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CL2024★ 6 cited

Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis

Shaochen Xu, Zihao Wu, Huaqin Zhao +7

In this study, we leverage LLM to enhance the semantic analysis and develop similarity metrics for texts, addressing the limitations of traditional unsupervised NLP metrics like RO…

eess.AS2023★ 7 cited

Exploring Multimodal Approaches for Alzheimer's Disease Detection Using Patient Speech Transcript and Audio Data

Hongmin Cai, Xiaoke Huang, Zhengliang Liu +8

Alzheimer's disease (AD) is a common form of dementia that severely impacts patient health. As AD impairs the patient's language understanding and expression ability, the speech of…

cs.CL2023★ 79 cited

Differentiate ChatGPT-generated and Human-written Medical Texts

Wenxiong Liao, Zhengliang Liu, Haixing Dai +8

Background: Large language models such as ChatGPT are capable of generating grammatically perfect and human-like text content, and a large number of ChatGPT-generated texts have ap…

cs.CL2023★ 99 cited

AugGPT: Leveraging ChatGPT for Text Data Augmentation

Haixing Dai, Zhengliang Liu, Wenxiong Liao +15

Text data augmentation is an effective strategy for overcoming the challenge of limited sample sizes in many natural language processing (NLP) tasks. This challenge is especially p…

cs.CL2023★ 45 cited

Mask-guided BERT for Few Shot Text Classification

Wenxiong Liao, Zhengliang Liu, Haixing Dai +11

Transformer-based language models have achieved significant success in various domains. However, the data-intensive nature of the transformer architecture requires much labeled dat…

cs.AI2022★ 4 cited

Coarse-to-fine Knowledge Graph Domain Adaptation based on Distantly-supervised Iterative Training

Hongmin Cai, Wenxiong Liao, Zhengliang Liu +12

Modern supervised learning neural network models require a large amount of manually labeled data, which makes the construction of domain-specific knowledge graphs time-consuming an…