From the 1 of 23 linked papers with an AI index.
5 papers · 1 filter
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Mengru Wang, Junfeng Fang, Shuofei Qiao +16
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI deve…
Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
Shuyi Miao, Wangjie Qiu, Pengyang Shao +4
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mech…
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Jinhe Bi, Chennan Zhou, Zengjie Jin +10
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories…
One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs
Enyi Shi, Fei Shen, Chuancheng Shi +4
The paper introduces a neuron‑level safety alignment method that identifies and updates a tiny set of shared safety neurons across languages and modalities, enabling large vision‑l…
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
Ruipeng Wang, Yuxin Chen, Yukai Wang +9
Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployment…