23 citations · 107 across the 30 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025★ 1 cited
Learning to Attribute with Attention
Benjamin Cohen-Wang, Yung-Sung Chuang, Aleksander Madry
Given a sequence of tokens generated by a language model, we may want to identify the preceding tokens that influence the model to generate this sequence. Performing such token att…
cs.LG2024★ 5 cited
Curiosity-driven Red-teaming for Large Language Models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang +5
Large language models (LLMs) hold great potential for many natural language applications but risk generating incorrect or toxic content. To probe when an LLM generates unwanted con…