activity
20202026
most citedMEDITRON-70B: Scaling Medical Pretraining for Large Language Models

123 citations · 123 across the 4 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut

Latent language identification is often used to argue that multilingual language models route computation through language-specific states, such as English pivots. However, existin…

cs.CL2025

Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining

Deniz Bayazit, Aaron Mueller, Antoine Bosselut

Large language models (LLMs) learn non-trivial abstractions during pretraining, such as detecting irregular plural noun subjects. However, because traditional evaluation methods (e…

cs.CL2023123 cited

MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Zeming Chen, Alejandro Hernández Cano, Angelika Romanou +17

Large language models (LLMs) can potentially democratize access to medical knowledge. While many efforts have been made to harness and improve LLMs' medical knowledge and reasoning…

cs.CL2023

Discovering Knowledge-Critical Subnetworks in Pretrained Language Models

Deniz Bayazit, Negar Foroutan, Zeming Chen +2

Pretrained language models (LMs) encode implicit representations of knowledge in their parameters. However, localizing these representations and disentangling them from each other…

cs.CL2023

PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives

Silin Gao, Beatriz Borges, Soyoung Oh +5

Sustaining coherent and engaging narratives requires dialogue or storytelling agents to understand how the personas of speakers or listeners ground the narrative. Specifically, the…