4 citations · 4 across the 6 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.CR2025
Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
Tianhang Zhao, Haodong Zhao, Wei Du +5
The ``Pre-train, then fine-tune'' paradigm has revolutionized Natural Language Processing (NLP). In this context, transferable backdoors pose a severe threat to the Pre-trained Lan…
cs.AI2025
NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
Yulong Zhang, Li Wang, Wei Du +7
Verifying multi-step reasoning in large language models is difficult due to imprecise error localization and high token costs. Existing methods either assess entire reasoning chain…