6 papers
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
Philipp D. Siedler, Jordan Sassoon
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of i…
Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation
Philipp D. Siedler
Strategic interaction in adversarial domains such as law, diplomacy, and negotiation is mediated by language, yet most game-theoretic models abstract away the mechanisms of persuas…
LLM-Mediated Guidance of MARL Systems
Philipp D. Siedler, Ian Gemp
In complex multi-agent environments, achieving efficient learning and desirable behaviours is a significant challenge for Multi-Agent Reinforcement Learning (MARL) systems. This wo…
SPhyR: Spatial-Physical Reasoning Benchmark on Material Distribution
Philipp D. Siedler
We introduce a novel dataset designed to benchmark the physical and spatial reasoning capabilities of Large Language Models (LLM) based on topology optimization, a method for compu…
HIVEX: A High-Impact Environment Suite for Multi-Agent Research (extended version)
Philipp Dominic Siedler
Games have been vital test beds for the rapid development of Agent-based research. Remarkable progress has been achieved in the past, but it is unclear if the findings equip for re…
Learning to Communicate and Collaborate in a Competitive Multi-Agent Setup to Clean the Ocean from Macroplastics
Philipp Dominic Siedler
Finding a balance between collaboration and competition is crucial for artificial agents in many real-world applications. We investigate this using a Multi-Agent Reinforcement Lear…