2 papers
cs.AI2026
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
Haoyue Yang, Zhangxiao Shen, Fan Ding +8
Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human…
cs.SE2024
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
Linyi Li, Shijie Geng, Zhenwen Li +7
Large Language Models for code (code LLMs) have witnessed tremendous progress in recent years. With the rapid development of code LLMs, many popular evaluation benchmarks, such as…