5 papers
Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines
Tanzim Ahad, Ismail Hossain, Md Jahangir Alam +4
We introduce Semantic Intent Fragmentation (SIF), an attack class against LLM orchestration systems where a single, legitimately phrased request causes an orchestrator to decompose…
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
Ismail Hossain, Sai Puppala, Jannatul Ferdaus +4
A guard model fine-tuned on entirely benign data can lose all safety alignment -- not through adversarial manipulation, but through standard domain specialization. We demonstrate t…
Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents
Sai Puppala, Ismail Hossain, Md Jahangir Alam +5
Large language models are increasingly deployed as *deep agents* that plan, maintain persistent state, and invoke external tools, shifting safety failures from unsafe text to unsaf…
ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
Yoonpyo Lee
Designing nuclear reactor cores requires navigating large discrete design spaces governed by complex neutronic interactions. Traditional deterministic, metaheuristic, and machine-l…
Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
Yoon Pyo Lee
The integration of Large Language Models (LLMs) into safety-critical domains, such as nuclear engineering, necessitates a deep understanding of their internal reasoning processes.…