5 papers
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
Chenxu Liu, Yingjie Fu, Wei Yang +2
Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential. However, building a…
Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs
Yuren Mao, Peigen Liu, Xinjian Wang +11
Multi-Agent System (MAS) developing frameworks serve as the foundational infrastructure for social simulations powered by Large Language Models (LLMs). However, existing frameworks…
KnowThyself: An Agentic Assistant for LLM Interpretability
Suraj Prasai, Mengnan Du, Ying Zhang +1
We develop KnowThyself, an agentic assistant that advances large language model (LLM) interpretability. Existing tools provide useful insights but remain fragmented and code-intens…
Temac: Multi-Agent Collaboration for Automated Web GUI Testing
Chenxu Liu, Zhiyu Gu, Guoquan Wu +3
Quality assurance of web applications is critical, as web applications play an essential role in people's daily lives. To reduce labor costs, automated web GUI testing (AWGT) is wi…
Data and System Perspectives of Sustainable Artificial Intelligence
Tao Xie, David Harel, Dezhi Ran +11
Sustainable AI is a subfield of AI for concerning developing and using AI systems in ways of aiming to reduce environmental impact and achieve sustainability. Sustainable AI is inc…