From the 1 of 15 linked papers with an AI index.
15 papers
Memorization Diagnostics for Code LLMs Should be Scale-Aware
Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré +6
The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread me…
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2
The paper evaluates three commercial image‑moderation services built on foundation models and shows that simple, model‑agnostic image transformations (e.g., color inversion, graysc…
Predictable Confabulations: Factual Recall by LLMs Scales with Model Size and Topic Frequency
Matthew L. Smith, Jonathan P. Shock, Samuel T. Segun +2
While scaling laws govern aggregate large language model performance, no scaling law has linked factual recall to both model size and training-data composition. We evaluated 38 mod…
Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?
Prateek Rajput, Yewei Song, Iyiola E. Olatunji +2
Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? To investigate this, we induce p…
SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples
Dren Fazlija, Monty-Maximilian Zühlke, Johanna Schrader +4
Unrestricted adversarial attacks aim to fool computer vision models without being constrained by -norm bounds to remain imperceptible to humans, for example, by changing an…
From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software
Moustapha Awwalou Diouf, Maimouna Tamah Diao, Iyiola Emmanuel Olatunji +6
LLMs democratize software engineering by enabling non-programmers to create applications, but this same accessibility fundamentally undermines security assumptions that have guided…