Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
Xu Wang, Yan Hu, Benyou Wang +1
Sparse Autoencoders (SAEs) are widely used to steer large language models (LLMs), based on the assumption that their interpretable features naturally enable effective model behavio…
cs.LG2025
Federated Linear Dueling Bandits
Xuhan Huang, Yan Hu, Zhiyan Li +3
Contextual linear dueling bandits have recently garnered significant attention due to their widespread applications in important domains such as recommender systems and large langu…