◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Wei Shi

3 papers hereh-index 475 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2

Across the 2 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.CL1
same name
  • Wei Shi — 21 papers
  • Wei Shi — 8 papers
  • Wei Shi — 6 papers, h 19
  • Wei Shi — 6 papers, h 3
  • Wei Shi — 5 papers
  • Wei Shi — 5 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

collaborators

3 papers

cs.LG2025

Interpretable Reward Model via Sparse Autoencoder

Shuyi Zhang, Wei Shi, Sihang Li +3

Large language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human…

cs.CL2025

SAFER: Probing Safety in Reward Models with Sparse Autoencoder

Wei Shi, Ziyuan Xie, Sihang Li +1

Reinforcement learning from human feedback (RLHF) is a key paradigm for aligning large language models (LLMs) with human values, yet the reward models at its core remain largely op…

cs.LG2025

Route Sparse Autoencoder to Interpret Large Language Models

Wei Shi, Sihang Li, Tao Liang +4

Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning. Sparse autoencoders (SAEs) have demons…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.