2 papers
cs.CV2026
Can a Teenager Fool an AI? Evaluating Low-Cost Cosmetic Attacks on Age Estimation Systems
Xingyu Shen, Tommy Duong, Xiaodong An +6
Age estimation systems are increasingly deployed as gatekeepers for age-restricted online content, yet their robustness to cosmetic modifications has not been systematically evalua…
cs.CL2025
Counting Clues: A Lightweight Probabilistic Baseline Can Match an LLM
Furong Jia, Yuan Pu, Finn Guo +1
Large language models (LLMs) excel on multiple-choice clinical diagnosis benchmarks, yet it is unclear how much of this performance reflects underlying probabilistic reasoning. We…