1 citations · 1 across the 8 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Towards Context-Invariant Safety Alignment for Large Language Models
Yixu Wang, Yang Yao, Xin Wang +4
Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a standard prompt but comply when t…
cs.CL2025
Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
Yixin Cao, Shibo Hong, Xinze Li +24
Large Language Models (LLMs) are advancing at an amazing speed and have become indispensable across academia, industry, and daily applications. To keep pace with the status quo, th…
cs.CL2025
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
Hongyu Su, Yifeng Gao, Yifan Ding +1
The rapid advancement of Large Language Models (LLMs) has increased the complexity and cost of fine-tuning, leading to the adoption of API-based fine-tuning as a simpler and more e…