24 citations · 24 across the 2 of their papers we have counts for
2 papers
cs.CL2026
Beyond Shallow Alignment: How Post-Training Methods Determine Refusal Circuits And Steering Robustness
Hoang Cuong Nguyen, Mark Dras, Usman Naseem
How do the methods used to train language models to refuse harmful requests shape how that refusal actually works inside the model? We compare three post-training methods - supervi…
cs.CR2025★ 24 cited
Towards Effective Identification of Attack Techniques in Cyber Threat Intelligence Reports using Large Language Models
Hoang Cuong Nguyen, Shahroz Tariq, Mohan Baruwal Chhetri +1
This work evaluates the performance of Cyber Threat Intelligence (CTI) extraction methods in identifying attack techniques from threat reports available on the web using the MITRE…