123 citations · 123 across the 4 of their papers we have counts for
5 papers · 1 filter
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut
Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existin…
Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
Deniz Bayazit, Aaron Mueller, Antoine Bosselut
Large language models (LLMs) learn non-trivial abstractions during pretraining, such as detecting irregular plural noun subjects. However, because traditional evaluation methods (e…
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
Zeming Chen, Alejandro Hernández Cano, Angelika Romanou +17
Large language models (LLMs) can potentially democratize access to medical knowledge. While many efforts have been made to harness and improve LLMs' medical knowledge and reasoning…
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
Deniz Bayazit, Negar Foroutan, Zeming Chen +2
Pretrained language models (LMs) encode implicit representations of knowledge in their parameters. However, localizing these representations and disentangling them from each other…
PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives
Silin Gao, Beatriz Borges, Soyoung Oh +5
Sustaining coherent and engaging narratives requires dialogue or storytelling agents to understand how the personas of speakers or listeners ground the narrative. Specifically, the…