2 papers
cs.AI2026
GAVEL: Towards Rule-Based Safety Through Activation Monitoring
Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower +3
Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. Ho…
cs.CR2026★ 1 cited
Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
Gilad Gressel, Rahul Pankajakshan, Shir Rozenfeld +4
Romance-baiting scams have become a major source of financial and emotional harm worldwide. These operations are run by organized crime syndicates that traffic thousands of people…