Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Challenges in Ensuring AI Safety in DeepSeek-R1 Models: The Shortcomings of Reinforcement Learning Strategies
Manojkumar Parmar, Yuvaraj Govindarajulu
Large Language Models (LLMs) have achieved remarkable progress in reasoning, alignment, and task-specific performance. However, ensuring harmlessness in these systems remains a cri…
cs.LG2024
MISLEAD: Manipulating Importance of Selected features for Learning Epsilon in Evasion Attack Deception
Vidit Khazanchi, Pavan Kulkarni, Yuvaraj Govindarajulu +1
Emerging vulnerabilities in machine learning (ML) models due to adversarial attacks raise concerns about their reliability. Specifically, evasion attacks manipulate models by intro…