From the 1 of 23 linked papers with an AI index.
23 papers
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
Mengru Wang, Junfeng Fang, Shuofei Qiao +16
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI deve…
Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
Shuyi Miao, Wangjie Qiu, Pengyang Shao +4
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mech…
Continual Learning in Transition
Zhiyan Hou, Dan Zhang, Tao Feng +11
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architect…
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Aniri, Jinhe Bi, Peng Liao +5
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw pr…
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Jinhe Bi, Chennan Zhou, Zengjie Jin +10
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories…
SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs
Jian Yu, Fei Shen, Cong Wang +6
Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes a significant gap between e…