10 papers
Moravec's Paradox: Towards an Auditory Turing Test
David Noever, Forrest McKee
This research work demonstrates that current AI systems fail catastrophically on auditory tasks that humans perform effortlessly. Drawing inspiration from Moravec's paradox (i.e.,…
Favicon Trojans: Executable Steganography Via Ico Alpha Channel Exploitation
David Noever, Forrest McKee
This paper presents a novel method of executable steganography using the alpha transparency layer of ICO image files to embed and deliver self-decompressing JavaScript payloads wit…
Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
David Noever, Forrest McKee
This study explores Large Language Models (LLMs) as autonomous agents for real-world tasks, including freelance software development. This work presents a new benchmark that evalua…
Alpha Excel Benchmark
David Noever, Forrest McKee
This study presents a novel benchmark for evaluating Large Language Models (LLMs) using challenges derived from the Financial Modeling World Cup (FMWC) Excel competitions. We intro…
Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests
David Noever, Forrest McKee
The development of robust safety benchmarks for large language models requires open, reproducible datasets that can measure both appropriate refusal of harmful content and potentia…
Dueling QR Codes: The Hyding of Dr. Jeckyl
David Noever, Forrest McKee
The paper presents a novel technique for encoding dual messages within standard Quick Response (QR) codes through precise half-pixel module splitting. This work challenges fundamen…