activity
20242026
most citedTSPRank: Bridging Pairwise and Listwise Methods with a Bilinear Travelling Salesman Model

3 citations · 5 across the 20 of their papers we have counts for

collaborators
Showing cs.CLShow all

14 papers · 1 filter

cs.CL2026

Debiasing Without Protected Attributes: Latent Concept Erasure from Textual Profiles

Shun Shao, Zheng Zhao, Anna Korhonen +2

Most fairness research in NLP assumes direct access to protected attributes such as gender, race, or nationality. In practice, however, such information is often unavailable due to…

cs.CL2026

Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer

Shun Shao, Binxu Wang, Shay B. Cohen +2

Mechanistic interpretability has made it possible to localize circuits underlying specific behaviors in language models, but existing methods are expensive, model-specific, and dif…

cs.CL2026

Self Knowledge Re-expression: A Fully Local Method for Adapting LLMs to Tasks Using Intrinsic Knowledge

Mengyu Wang, Xiaoying Zhi, Zhiyi Li +4

While the next-token prediction (NTP) paradigm enables large language models (LLMs) to express their intrinsic knowledge, its sequential nature constrains performance on specialize…

cs.CL2026

MoRFI: Monotonic Sparse Autoencoder Feature Identification

Dimitris Dimakopoulos, Shay B. Cohen, Ioannis Konstas

Large language models (LLMs) acquire most of their factual knowledge during the pre-training stage, through next token prediction. Subsequent stages of post-training often introduc…

cs.CL2026

Old Habits Die Hard: How Conversational History Geometrically Traps LLMs

Adi Simhi, Fazl Barez, Martin Tutek +2

How does the conversational past of large language models (LLMs) influence their future performance? Recent work suggests that LLMs are affected by their conversational history in…

cs.CL2026

Spectral Attention Steering for Prompt Highlighting

Weixian Waylon Li, Yuchen Niu, Yongxin Yang +3

Attention steering is an important technique for controlling model focus, enabling capabilities such as prompt highlighting, where the model prioritises user-specified text. Howeve…