5 papers
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Dahlia Shehata, Ming Li
We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index (), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic…
The Bystander Effect in Multi-Agent Reasoning: Quantifying Cognitive Loafing in Collaborative Interactions
Dahlia Shehata, Ming Li
Multi-agent systems (MAS) assume that collaborating inherently improves Large Language Model (LLM) reasoning. We challenge this by demonstrating that simulated social pressure trig…
The Inverse-Wisdom Law: Architectural Tribalism and the Consensus Paradox in Agentic Swarms
Dahlia Shehata, Ming Li
As AI transitions toward multi-agent systems (MAS) to solve complex workflows, research paradigms operate on the axiomatic assumption that agent collaboration mirrors the "Wisdom o…
Beyond the Attention Stability Boundary: Agentic Self-Synthesizing Reasoning Protocols
Dahlia Shehata, Ming Li
As LLM agents transition to autonomous digital coworkers, maintaining deterministic goal-directedness in non-linear multi-turn conversations emerged as an architectural bottleneck.…
Rumour Evaluation with Very Large Language Models
Dahlia Shehata, Robin Cohen, Charles Clarke
Conversational prompt-engineering-based large language models (LLMs) have enabled targeted control over the output creation, enhancing versatility, adaptability and adhoc retrieval…