2 papers
cs.CL2026
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle +10
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, d…
cs.LG2025
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
Femi Bello, Anubrata Das, Fanzhi Zeng +2
It has been hypothesized that neural networks with similar architectures trained on similar data learn shared representations relevant to the learning task. We build on this idea b…