Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
arXiv:2606.28345 · doi:10.1145/3805689.3812366
Abstract
LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by age, status, and group size, failure to calibrate pluralistically can scale into unequal access. Yet LLM moral audits remain English-centered, rarely test embodied contexts, leaving pluralistic calibration as an urgent diagnostic gap amid intensifying LLM-robot deployment. We introduce a gradient-based audit framework for multilingual evaluation of LLM moral trade-off behavior against cultural preference gradients. Grounded in nine cross-domain social robotics reviews (>8,000 papers), we derive symmetry-controlled scenarios across care, education, and services, translating the Moral Machine Experiment's "whom to spare" into "whom to assist first" dilemmas with preserved identity trade-offs (many vs. few; young vs. old; higher vs. lower status). We audit four LLMs across four country-language pairs in four prompting regimes (57,600 decisions), benchmarked against country-specific MME preference gradients. Ordinal concordance tests whether models differentiate cultural contexts; a governance typology maps vulnerabilities in gradient differentiation, directional tendency, and deliberation. We find persistent, culturally asymmetric gradient tracking failures that prompting alone cannot reliably correct: quality calibration is nearly twice as strong for Western-language decisions as for Chinese and Japanese; high determinism in majority-first trade-offs often erases cross-cultural gradients; partial sensitivity to age- and status-based norms risks sidelining minorities. Prompting effects are uneven; only contrastive exemplars yield consistent gains, while reasoning-only prompts can worsen tracking. Our results motivate multilingual, pluralistic audits as an LLM-robot pre-deployment gate and suggest model factors are a more robust lever than prompting alone.
Accepted for publication in Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)
References in corpus (19)
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- PaLM-E: An Embodied Multimodal Language Model
- Symmetry reduction for testing -block-positivity via extendibility
- WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT?
- A Bibliometric Review of Large Language Models Research from 2017 to 2023
- The Moral Machine Experiment on Large Language Models
- Large Language Models for Robotics: A Survey
- From Learning to Relearning: A Framework for Diminishing Bias in Social Robot Navigation
- Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models
- What are human values, and how do we align AI to them?
- Towards Measuring and Modeling "Culture" in LLMs: A Survey
- Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
- Probing the Moral Development of Large Language Models through Defining Issues Test
- Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
- Emergent Abilities in Large Language Models: A Survey
- From Text to Motion: Grounding GPT-4 in a Humanoid Robot "Alter3"
- Does Moral Code Have a Moral Code? Probing Delphi's Moral Philosophy
- Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment