activity
20232026
most citedAI Alignment: A Comprehensive Survey

76 citations · 76 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL2026

PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users

Hannah Rose Kirk, Liu Leqi, Fanzhi Zeng +4

Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simu…

cs.LG2026

Curriculum Learning-Guided Progressive Distillation in Large Language Models

Jincheng Cao, Fanzhi Zeng, Leqi Liu +1

Knowledge distillation is a key technique for transferring the capabilities of large language models (LLMs) into smaller, more efficient student models. Existing distillation appro…

cs.AI2025

ADRD: LLM-Driven Autonomous Driving Based on Rule-based Decision Systems

Fanzhi Zeng, Siqi Wang, Chuzhao Zhu +1

How to construct an interpretable autonomous driving decision-making system has become a focal point in academic research. In this study, we propose a novel approach that leverages…

cs.LG2025

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models

Femi Bello, Anubrata Das, Fanzhi Zeng +2

It has been hypothesized that neural networks with similar architectures trained on similar data learn shared representations relevant to the learning task. We build on this idea b…

cs.LG2024

Reward Generalization in RLHF: A Topological Perspective

Tianyi Qiu, Fanzhi Zeng, Jiaming Ji +7

Existing alignment methods share a common topology of information flow, where reward information is collected from humans, modeled with preference learning, and used to tune langua…

cs.AI2023★ 76 cited

AI Alignment: A Comprehensive Survey

Jiaming Ji, Tianyi Qiu, Boyuan Chen +23

AI alignment aims to make AI systems behave in line with human intentions and values. As AI systems grow more capable, so do risks from misalignment. To provide a comprehensive and…