Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
CogToM: A Comprehensive Theory of Mind Benchmark inspired by Human Cognition for Large Language Models
Haibo Tong, Zeyang Yue, Feifei Zhao +6
Whether Large Language Models (LLMs) truly possess human-like Theory of Mind (ToM) capabilities has garnered increasing attention. However, existing benchmarks remain largely restr…
cs.AI2025
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
Ruolin Chen, Yinqian Sun, Jihang Wang +3
Embodied agents powered by large language models (LLMs) inherit advanced planning capabilities; however, their direct interaction with the physical world exposes them to safety vul…