3 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CL2024
Model Attribution in LLM-Generated Disinformation: A Domain Generalization Approach with Supervised Contrastive Learning
Alimohammad Beigi, Zhen Tan, Nivedh Mudiam +3
Model attribution for LLM-generated disinformation poses a significant challenge in understanding its origins and mitigating its spread. This task is especially challenging because…
cs.CL2023★ 3 cited
Interpreting Pretrained Language Models via Concept Bottlenecks
Zhen Tan, Lu Cheng, Song Wang +3
Pretrained language models (PLMs) have made significant strides in various natural language processing tasks. However, the lack of interpretability due to their ``black-box'' natur…