works on

From the 1 of 23 linked papers with an AI index.

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

Mengru Wang, Junfeng Fang, Shuofei Qiao +16

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI deve…

cs.AI2026

Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways

Shuyi Miao, Wangjie Qiu, Pengyang Shao +4

Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mech…

cs.AI2026

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Jinhe Bi, Chennan Zhou, Zengjie Jin +10

On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories…

cs.AI2026

One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs

Enyi Shi, Fei Shen, Chuancheng Shi +4

The paper introduces a neuron‑level safety alignment method that identifies and updates a tiny set of shared safety neurons across languages and modalities, enabling large vision‑l…

cs.AI2026

AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition

Ruipeng Wang, Yuxin Chen, Yukai Wang +9

Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployment…