4 papers · 1 filter
From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing
Wei Liu, Hongkai Liu, Zhiying Deng +2
LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the target vector to multiple pr…
The Cylindrical Representation Hypothesis for Language Model Steering
Lang Gao, Jinghui Zhang, Wei Liu +7
Steering is a widely used technique for controlling large language models, yet its effects are often unstable and hard to predict. Existing theoretical accounts are largely based o…
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
Lang Gao, Xuhui Li, Chenxi Wang +7
Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this ability also heightens the risk of…
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
Lang Gao, Kaiyang Wan, Wei Liu +6
Bias in Large Language Models (LLMs) significantly undermines their reliability and fairness. We focus on a common form of bias: when two reference concepts in the model's concept…