From the 1 of 18 linked papers with an AI index.
18 papers
Memorization Diagnostics for Code LLMs Should be Scale-Aware
Prateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djiré +6
The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current literature frequently reports widespread me…
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
PaweŠBorsukiewicz, Daniele Lunghi, Wendkûuni C. Ouédraogo +2
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are tr…
Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2
The paper evaluates three commercial image‑moderation services built on foundation models and shows that simple, model‑agnostic image transformations (e.g., color inversion, graysc…
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications
Despoina Giarimpampa, Roland Meier, Tegawendé F. Bissyandé +2
Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SO…
A Single Patch Is Not Enough: Deterministic Fusion of Repair Candidates
Boyang Yang, Xiangliang Hu, Luyao Ren +4
Modern LLM coding agents are commonly evaluated using pass@k, but developers typically apply a single final patch in real-world settings. This pass@k-to-pass@1 gap is a post-genera…
Detecting Malicious Agent Skills in the Wild using Attention
Bacem Etteib, Daniele Lunghi, Tégawendé F. Bissyandé
LLM agents increasingly load skills, file-based packages of natural-language instructions written by third parties and distributed through marketplaces, that execute with the user'…