Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift
San Kim, JinYeong Bak
Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially rele…
cs.CL2026
Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language Models
San Kim, Gary Geunbae Lee
Large Language Models (LLMs) have greatly advanced Natural Language Processing (NLP), particularly through instruction tuning, which enables broad task generalization without addit…
cs.CL2025
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
San Kim, Jonghwi Kim, Yejin Jeon +1
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by providing external knowledge for accurate and up-to-date responses. However, this reliance on external…