From the 1 of 17 linked papers with an AI index.
1 citations · 1 across the 9 of their papers we have counts for
6 papers · 1 filter
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang +23
The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…
Benchmarking at the Edge of Comprehension
Samuele Marro, Jialin Yu, Emanuele La Malfa +8
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improv…
End-to-end PDDL Planning with Hardcoded and Dynamic Agents
Emanuele La Malfa, Ping Zhu, Samuele Marro +2
We present an end-to-end framework for planning supported by verifiers. An orchestrator receives a human specification written in natural language and converts it into a PDDL (Plan…
Profit is the Red Team: Stress-Testing Agents in Strategic Economic Interactions
Shouqiao Wang, Marcello Politi, Samuele Marro +1
As agentic systems move into real-world deployments, their decisions increasingly depend on external inputs such as retrieved content, tool outputs, and information provided by oth…
A Scalable Communication Protocol for Networks of Large Language Models
Samuele Marro, Emanuele La Malfa, Jesse Wright +4
Communication is a prerequisite for collaboration. When scaling networks of AI-powered agents, communication must be versatile, efficient, and portable. These requisites, which we…
A Notion of Complexity for Theory of Mind via Discrete World Models
X. Angelo Huang, Emanuele La Malfa, Samuele Marro +3
Theory of Mind (ToM) can be used to assess the capabilities of Large Language Models (LLMs) in complex scenarios where social reasoning is required. While the research community ha…