2 citations · 3 across the 8 of their papers we have counts for
10 papers
Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks
Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah +1
Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cyberse…
PEA: Enhancing LLM Performance on Computational-Reasoning Tasks
Zi Wang, Shiwei Weng, Mohannad Alhanahnah +2
Large Language Models (LLMs) have exhibited remarkable capabilities across diverse domains, prompting investigations into their potential as generic reasoning engines. While recent…
SoK: Software Debloating Landscape and Future Directions
Mohannad Alhanahnah, Yazan Boshmaf, Ashish Gehani
Software debloating seeks to mitigate security risks and improve performance by eliminating unnecessary code. In recent years, a plethora of debloating tools have been developed, c…
DepsRAG: Towards Agentic Reasoning and Planning for Software Dependency Management
Mohannad Alhanahnah, Yazan Boshmaf
In the era of Large Language Models (LLMs) with their advanced capabilities, a unique opportunity arises to develop LLM-based digital assistant tools that can support software deve…
An Empirical Evaluation of Pre-trained Large Language Models for Repairing Declarative Formal Specifications
Mohannad Alhanahnah, Md Rashedul Hasan, Lisong Xu +1
Automatic Program Repair (APR) has garnered significant attention as a practical research domain focused on automatically fixing bugs in programs. While existing APR techniques pri…
slash: A Technique for Static Configuration-Logic Identification
Mohannad Alhanahnah, Philipp Schubert, Thomas Reps +2
Researchers have recently devised tools for debloating software and detecting configuration errors. Several of these tools rely on the observation that programs are composed of an…