Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
Jianhao Chen, Mayi Xu, Haoyang Chen +6
Large Reasoning Models (LRMs) face a distinct safety vulnerability: their internal reasoning chains may generate harmful content even when the final output appears benign. To addre…
cs.AI2025
Aligning VLM Assistants with Personalized Situated Cognition
Yongqi Li, Shen Zhou, Xiaohu Li +9
Vision-language models (VLMs) aligned with general human objectives, such as being harmless and hallucination-free, have become valuable assistants of humans in managing visual tas…