3 papers
cs.CR2026
PAL*M: Property Attestation for Large Generative Models
Prach Chantasantitam, Adam Ilyas Caulfield, Vasisht Duddu +2
Machine learning property attestations allow provers (e.g., model providers or owners) to attest properties of their models/datasets to verifiers (e.g., regulators, customers), ena…
cs.CR2026
Backdooring Bias in Large Language Models
Anudeep Das, Prach Chantasantitam, Gurjot Singh +3
Large language models (LLMs) are increasingly deployed in settings where inducing a bias toward a certain topic can have significant consequences, and backdoor attacks can be used…
cs.CV2025
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?
Anudeep Das, Gurjot Singh, Prach Chantasantitam +1
Generative models, particularly diffusion-based text-to-image (T2I) models, have demonstrated astounding success. However, aligning them to avoid generating content with unacceptab…