6 papers
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige +3
LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditi…
SADE: Symptom-Aware Diagnostic Escalation for LLM-Based Network Troubleshooting
Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige +3
Large language model (LLM) agents are increasingly applied to network troubleshooting, but root-cause localization on public benchmarks remains well below practical deployment thre…
Pen-Strategist: A Reasoning Framework for Penetration Testing Strategy Formation and Analysis
Yasod Ginige, Pasindu Marasinghe, Sajal Jain +1
Cyber threats are rapidly increasing, expanding their impact from large-scale enterprises to government services and individual users, making robust security systems increasingly e…
PrivPRISM: Automatically Detecting Discrepancies Between Google Play Data Safety Declarations and Developer Privacy Policies
Bhanuka Silva, Dishanika Denipitiyage, Anirban Mahanti +2
End-users seldom read verbose privacy policies, leading app stores like Google Play to mandate simplified data safety declarations as a user-friendly alternative. However, these se…
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
Yasod Ginige, Akila Niroshan, Sajal Jain +1
Penetration testing and vulnerability assessment are essential industry practices for safeguarding computer systems. As cyber threats grow in scale and complexity, the demand for p…
Detecting Content Rating Violations in Android Applications: A Vision-Language Approach
D. Denipitiyage, B. Silva, S. Seneviratne +2
Despite regulatory efforts to establish reliable content-rating guidelines for mobile apps, the process of assigning content ratings in the Google Play Store remains self-regulated…