works on

From the 2 of 10 linked papers with an AI index.

collaborators

10 papers

cs.SE2026

Agent-Based Software Artifact Evaluation

Zhaonan Wu, Yanjie Zhao, Zhenpeng Chen +2

The paper introduces ArtifactGuide, a structured scoring rubric for software artifact evaluation, and ArtifactCopilot, an LLM‑based agent that follows this rubric to automatically…

cs.CL2026

Ensemble Learning for Large Language Models in Text and Code Generation: A Survey

Mari Ashiga, Wei Jie, Fan Wu +5

Generative Pretrained Transformers (GPTs) are foundational Large Language Models (LLMs) for text generation. However, individual LLMs often produce inconsistent outputs and exhibit…

cs.LG2026

Building Production-Ready Probes For Gemini

János Kramár, Joshua Engels, Zheng Wang +4

Frontier language model capabilities are improving rapidly. We thus need stronger mitigations against bad actors misusing increasingly powerful systems. Prior work has shown that a…

cs.AI2026

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence

Huan-ang Gao, Jiayi Geng, Wenyue Hua +24

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks but remain fundamentally static, unable to adapt their internal parameters to novel task…

cs.SE2025

Evolving Excellence: Automated Optimization of LLM-based Agents

Paul Brookes, Vardan Voskanyan, Rafail Giavrimis +18

Agentic AI systems built on large language models (LLMs) offer significant potential for automating complex workflows, from software development to customer support. However, LLM a…

cs.AI2025

D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies

Sen Chen, Tong Zhao, Yi Bin +3

Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial Ge…