2 papers
cs.AI2026
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Zejun Xu, Taiyi Chen, Jin Li +13
Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate…
cs.CR2026
MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization
Yongtong Gu, Songze Li, Xia Hu
The increasing misuse of AI-generated texts (AIGT) has motivated the rapid development of AIGT detection methods. However, the reliability of these detectors remains fragile agains…