2 papers
cs.CR2026
aCAPTCHA: Verifying That an Entity Is a Capable Agent via Asymmetric Hardness
Zuyao Xu, Xiang Li, Fubin Wu +3
As autonomous AI agents increasingly populate the Internet, a novel security challenge arises: "Is this entity an AI agent?" It is a new entity-type verification problem with no es…
cs.CL2025
Evaluation of OpenAI o1: Opportunities and Challenges of AGI
Tianyang Zhong, Zhengliang Liu, Yi Pan +73
This comprehensive study evaluates the performance of OpenAI's o1-preview large language model across a diverse array of complex reasoning tasks, spanning multiple domains, includi…