2 papers
cs.AI2026
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
Haoyue Yang, Zhangxiao Shen, Fan Ding +8
Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human…
cs.CE2026
Are LLMs Socially Adaptive? Contrasting Belief Evolution in Large Language Models and Humans
Yu Lei, Hao Liu, Chengxing Xie +6
As large language models (LLMs) increasingly engage in complex social interactions, ensuring that their behaviors align with human ethical principles and intentions, known as value…