2 papers
cs.CL2026
Do Activation Verbalization Methods Convey Privileged Information?
Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers +2
Recent interpretability methods have proposed to translate LLM internal representations into natural language descriptions using a second verbalizer LLM. This is intended to illumi…
cs.LG2026
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
Hiba Ahsan, Byron C. Wallace
LLMs are increasingly being used in healthcare. This promises to free physicians from drudgery, enabling better care to be delivered at scale. But the use of LLMs in this space als…