6 papers
CRC-Screen: Certified DNA-Synthesis Hazard Screening Under Taxonomic Shift
Najmul Hasan
DNA-synthesis providers screen incoming orders by searching the requested sequence against curated hazard lists. We show that this baseline collapses to a 100% false-flag rate when…
GRPO Does Not Close the Multi-Agent Coordination Gap
Najmul Hasan, Prashanth BusiReddyGari
We measure how well current large language models coordinate as multiple agents sharing a common resource, using the dining philosophers problem as a clean test bed. Across 630 epi…
DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention
Najmul Hasan, Prashanth BusiReddyGari
We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure task-level success under a fixed pr…
Honeypot Protocol
Najmul Hasan
Trusted monitoring, the standard defense in AI control, is vulnerable to adaptive attacks, collusion, and strategic attack selection. All of these exploit the fact that monitoring…
Time-Complexity Characterization of NIST Lightweight Cryptography Finalists
Najmul Hasan, Prashanth BusiReddyGari
Lightweight cryptography is becoming essential as emerging technologies in digital identity systems and Internet of Things verification continue to demand strong cryptographic assu…
Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection
Najmul Hasan, Prashanth BusiReddyGari
The Uniform Resource Locator (URL), introduced in a connectivity-first era to define access and locate resources, remains historically limited, lacking future-proof mechanisms for…