2 papers
cs.CR2025
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
Sheng Yin, Xianghe Pang, Yuanzhuo Ding +7
With the integration of large language models (LLMs), embodied agents have strong capabilities to understand and plan complicated natural language instructions. However, a foreseea…
cs.CL2024
Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review
Rui Ye, Xianghe Pang, Jingyi Chai +6
Scholarly peer review is a cornerstone of scientific advancement, but the system is under strain due to increasing manuscript submissions and the labor-intensive nature of the proc…