4 papers
Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning
Jitao Sang, Yuhang Wang, Jing Zhang +5
This paper presents a follow-up study to OpenAI's recent superalignment work on Weak-to-Strong Generalization (W2SG). Superalignment focuses on ensuring that high-level AI systems…
You talk what you read: Understanding News Comment Behavior by Dispositional and Situational Attribution
Yuhang Wang, Yuxiang Zhang, Dongyuan Lu +1
Many news comment mining studies are based on the assumption that comment is explicitly linked to the corresponding news. In this paper, we observed that users' comments are also h…
Towards Alleviating the Object Bias in Prompt Tuning-based Factual Knowledge Extraction
Yuhang Wang, Dongyuan Lu, Chao Kong +1
Many works employed prompt tuning methods to automatically optimize prompt queries and extract the factual knowledge stored in Pretrained Language Models. In this paper, we observe…
Universal Backdoor Attacks Detection via Adaptive Adversarial Probe
Yuhang Wang, Huafeng Shi, Rui Min +5
Extensive evidence has demonstrated that deep neural networks (DNNs) are vulnerable to backdoor attacks, which motivates the development of backdoor attacks detection. Most detecti…