2 papers
cs.CL2026
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
Zhengyang Tang, Yi Zhang, Chenxin Li +18
When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may be avoided because the agent re…
cs.AI2026
PeopleSearchBench: Evaluating AI-Powered People Search Platforms with Criteria-Grounded Verification
Tianyu Shi, Wei Wang, Zequn Xie +10
AI-powered people search platforms are increasingly deployed for recruiting, sales prospecting, and professional networking, yet no standardized benchmark exists for their rigorous…