284 citations · 1.3k across the 33 of their papers we have counts for
64 papers
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
Zeming Chen, Alejandro Hernández Cano, Angelika Romanou +17
Large language models (LLMs) can potentially democratize access to medical knowledge. While many efforts have been made to harness and improve LLMs' medical knowledge and reasoning…
SKILL: Structured Knowledge Infusion for Large Language Models
Fedor Moiseev, Zhe Dong, Enrique Alfonseca +1
Large language models (LLMs) have demonstrated human-level performance on a vast spectrum of natural language tasks. However, it is largely unexplored whether they can better inter…
Data-heterogeneity-aware Mixing for Decentralized Learning
Yatin Dandi, Anastasia Koloskova, Martin Jaggi +1
Decentralized learning provides an effective framework to train machine learning models with data distributed over arbitrary communication graphs. However, most existing approaches…
Improving Generalization via Uncertainty Driven Perturbations
Matteo Pagliardini, Gilberto Manunza, Martin Jaggi +2
Recently Shah et al., 2020 pointed out the pitfalls of the simplicity bias - the tendency of gradient-based algorithms to learn simple models - which include the model's high sensi…
Characterizing & Finding Good Data Orderings for Fast Convergence of Sequential Gradient Methods
Amirkeivan Mohtashami, Sebastian Stich, Martin Jaggi
While SGD, which samples from the data with replacement is widely studied in theory, a variant called Random Reshuffling (RR) is more common in practice. RR iterates through random…
Optimal Model Averaging: Towards Personalized Collaborative Learning
Felix Grimberg, Mary-Anne Hartley, Sai P. Karimireddy +1
In federated learning, differences in the data or objectives between the participating nodes motivate approaches to train a personalized machine learning model for each node. One s…