works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
most citedLLM Agents Are the Antidote to Walled Gardens

1 citations · 1 across the 9 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

Xiangning Lin, Shenzhe Zhu, Shu Yang +23

The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…

cs.AI2026

Benchmarking at the Edge of Comprehension

Samuele Marro, Jialin Yu, Emanuele La Malfa +8

As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improv…

cs.AI2026

End-to-end PDDL Planning with Hardcoded and Dynamic Agents

Emanuele La Malfa, Ping Zhu, Samuele Marro +2

We present an end-to-end framework for planning supported by verifiers. An orchestrator receives a human specification written in natural language and converts it into a PDDL (Plan…

cs.AI2026

Profit is the Red Team: Stress-Testing Agents in Strategic Economic Interactions

Shouqiao Wang, Marcello Politi, Samuele Marro +1

As agentic systems move into real-world deployments, their decisions increasingly depend on external inputs such as retrieved content, tool outputs, and information provided by oth…

cs.AI2024

A Scalable Communication Protocol for Networks of Large Language Models

Samuele Marro, Emanuele La Malfa, Jesse Wright +4

Communication is a prerequisite for collaboration. When scaling networks of AI-powered agents, communication must be versatile, efficient, and portable. These requisites, which we…

cs.AI2024

A Notion of Complexity for Theory of Mind via Discrete World Models

X. Angelo Huang, Emanuele La Malfa, Samuele Marro +3

Theory of Mind (ToM) can be used to assess the capabilities of Large Language Models (LLMs) in complex scenarios where social reasoning is required. While the research community ha…