fine-tuning 1large language models 1risk detection 1safety assessment 1semantic analysis 1subspace alignment 1
From the 1 of 6 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Kaicheng Shen, Lingyu Li, Wen Wu +3
AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks that may accumulate over ti…
cs.AI2026
MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
Liang Shan, Kaicheng Shen, Wen Wu +9
Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address implicit, domain-specific risks. T…