5 papers
SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents
Huaigang Yang, Ya Li, Min Ren +3
The paper introduces SAFERELBENCH, a benchmark that evaluates how vision‑language‑driven embodied agents maintain safety during actions by focusing on spatial relations like suppor…
NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms
Shiyun Zhao, Xinwei Song, Tianyu Guo +7
The paper presents NormAct, a benchmark for evaluating whether embodied planners using multimodal large language models can infer and follow hidden social norms while completing ta…
TongSIM: A General Platform for Simulating Intelligent Machines
Zhe Sun, Kunlun Wu, Chuanjian Fu +24
As artificial intelligence (AI) rapidly advances, especially in multimodal large language models (MLLMs), research focus is shifting from single-modality text processing to the mor…
Reasoning with Exploration: An Entropy Perspective
Daixuan Cheng, Shaohan Huang, Xuekai Zhu +4
Balancing exploration and exploitation is a central goal in reinforcement learning (RL). Despite recent advances in enhancing large language model (LLM) reasoning, most methods lea…
On Domain-Adaptive Post-Training for Multimodal Large Language Models
Daixuan Cheng, Shaohan Huang, Ziyu Zhu +5
Adapting general multimodal large language models (MLLMs) to specific domains, such as scientific and industrial fields, is highly significant in promoting their practical applicat…