2 papers
cs.CV2026
Concept Removal Guidance: Evidence-Calibrated Negative Guidance for Safe Diffusion Sampling
Yoonseok Choi, Chaeyoung Oh, Hyunjun Choi +2
Text-to-image diffusion models remain vulnerable to adversarial prompts that elicit disallowed content, motivating reliable inference-time controls. A popular approach is negative…
cs.AI2025
Monet: Mixture of Monosemantic Experts for Transformers
Jungwoo Park, Young Jin Ahn, Kee-Eung Kim +1
Understanding the internal computations of large language models (LLMs) is crucial for aligning them with human values and preventing undesirable behaviors like toxic content gener…