2 citations · 6 across the 13 of their papers we have counts for
5 papers · 1 filter
Learned or Memorized ? Quantifying Memorization Advantage in Code LLMs
Djiré Albérick Euraste, Kaboré Abdoul Kader, Jordan Samhi +3
The lack of transparency about code datasets used to train large language models (LLMs) makes it difficult to detect, evaluate, and mitigate data leakage. We present a perturbation…
Exploring Hidden Geographic Disparities in Android Apps
M. Alecci, P. Jiménez, J. Samhi +2
While mobile app evolution has been widely studied, geographical variation in app behavior remains largely unexplored. This paper presents a large-scale study of location-based And…
SIEVE: Towards Verifiable Certification for Code-datasets
Fatou Ndiaye Mbodji, El-hacen Diallo, Jordan Samhi +3
Code agents and empirical software engineering rely on public code datasets, yet these datasets lack verifiable quality guarantees. Static 'dataset cards' inform, but they are neit…
Beyond Language Barriers: Multi-Agent Coordination for Multi-Language Code Generation
Micheline Bénédicte Moumoula, Serge Lionel Nikiema, Albérick Euraste Djire +3
Producing high-quality code across multiple programming languages is increasingly important as today's software systems are built on heterogeneous stacks. Large language models (LL…
The Code Barrier: What LLMs Actually Understand?
Serge Lionel Nikiema, Jordan Samhi, Abdoul Kader Kaboré +2
Understanding code represents a core ability needed for automating software development tasks. While foundation models like LLMs show impressive results across many software engine…