6 citations · 6 across the 7 of their papers we have counts for
4 papers · 1 filter
MELT: A Behavioral Trace Dataset for High-Risk Memecoin Launch Detection
Sihao Hu, Selim Furkan Tekin, Yichang Xu +1
Launchpads have become the dominant mechanism for issuing memecoins, exposing investors to a new class of high-risk launches that existing rug-pull detection methods cannot capture…
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
Tiansheng Huang, Sihao Hu, Fatih Ilhan +2
Recent research demonstrates that the nascent fine-tuning-as-a-service business model exposes serious safety concerns: fine-tuning with a few harmful data uploaded from the users c…
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
Tiansheng Huang, Sihao Hu, Fatih Ilhan +4
Safety alignment is an important procedure before the official deployment of a Large Language Model (LLM). While safety alignment has been extensively studied for LLM, there is sti…
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
Tiansheng Huang, Sihao Hu, Fatih Ilhan +2
Recent research shows that Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks -- models lose their safety alignment ability after fine-tuning on a few harmf…