9 citations · 10 across the 12 of their papers we have counts for
6 papers · 1 filter
Mirror: A Multi-Agent System for AI-Assisted Ethics Review
Yifan Ding, Yuhui Shi, Zhiyan Li +10
Ethics review is a foundational mechanism of modern research governance, yet contemporary systems face increasing strain as ethical risks arise as structural consequences of large-…
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
Weiyi Wang, Xinchi Chen, Jingjing Gong +2
Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent be…
VehicleWorld: A Highly Integrated Multi-Device Environment for Intelligent Vehicle Interaction
Jie Yang, Jiajun Chen, Zhangyue Yin +7
Intelligent vehicle cockpits present unique challenges for API Agents, requiring coordination across tightly-coupled subsystems that exceed typical task environments' complexity. T…
FamilyTool: A Multi-hop Personalized Tool Use Benchmark
Yuxin Wang, Yiran Guo, Yining Zheng +7
The integration of tool learning with Large Language Models (LLMs) has expanded their capabilities in handling complex tasks by leveraging external tools. However, existing benchma…
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
Siyin Wang, Xingsong Ye, Qinyuan Cheng +5
As Artificial General Intelligence (AGI) becomes increasingly integrated into various facets of human life, ensuring the safety and ethical alignment of such systems is paramount.…
Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
Zhiyuan Zeng, Qinyuan Cheng, Zhangyue Yin +6
OpenAI o1 represents a significant milestone in Artificial Inteiligence, which achieves expert-level performances on many challanging tasks that require strong reasoning ability.Op…