From the 2 of 10 linked papers with an AI index.
10 papers
Agent-Based Software Artifact Evaluation
Zhaonan Wu, Yanjie Zhao, Zhenpeng Chen +2
The paper introduces ArtifactGuide, a structured scoring rubric for software artifact evaluation, and ArtifactCopilot, an LLM‑based agent that follows this rubric to automatically…
Ensemble Learning for Large Language Models in Text and Code Generation: A Survey
Mari Ashiga, Wei Jie, Fan Wu +5
Generative Pretrained Transformers (GPTs) are foundational Large Language Models (LLMs) for text generation. However, individual LLMs often produce inconsistent outputs and exhibit…
Building Production-Ready Probes For Gemini
János Kramár, Joshua Engels, Zheng Wang +4
Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
Huan-ang Gao, Jiayi Geng, Wenyue Hua +24
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks but remain fundamentally static, unable to adapt their internal parameters to novel task…
Evolving Excellence: Automated Optimization of LLM-based Agents
Paul Brookes, Vardan Voskanyan, Rafail Giavrimis +18
Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM a…
D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
Sen Chen, Tong Zhao, Yi Bin +3
Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial Ge…