activity
20192026
most citedOrca 2: Teaching Small Language Models How to Reason

32 citations · 51 across the 7 of their papers we have counts for

collaborators

8 papers

cs.SE2026★ 1 cited

The Llama 4 Herd: Architecture, Training, Evaluation, and Deployment Notes

Redacted by arXiv

This document consolidates publicly reported technical details about Metas Llama 4 model family. It summarizes (i) released variants (Scout and Maverick) and the broader herd conte…

cs.AI2023★ 32 cited

Orca 2: Teaching Small Language Models How to Reason

Arindam Mitra, Luciano Del Corro, Shweti Mahajan +12

Orca 1 learns from rich signals, such as explanation traces, allowing it to outperform conventional instruction-tuned models on benchmarks like BigBench Hard and AGIEval. In Orca 2…

cs.CL2023

MIReAD: Simple Method for Learning High-quality Representations from Scientific Documents

Anastasia Razdaibiedina, Alexander Brechalov

Learning semantically meaningful representations from scientific documents can facilitate academic literature search and improve performance of recommendation systems. Pre-trained…

cs.CL2023

Residual Prompt Tuning: Improving Prompt Tuning with Residual Reparameterization

Anastasia Razdaibiedina, Yuning Mao, Rui Hou +4

Prompt tuning is one of the successful approaches for parameter-efficient tuning of pre-trained language models. Despite being arguably the most parameter-efficient (tuned soft pro…

cs.CL2023★ 15 cited

Progressive Prompts: Continual Learning for Language Models

Anastasia Razdaibiedina, Yuning Mao, Rui Hou +3

We introduce Progressive Prompts - a simple and efficient approach for continual learning in language models. Our method allows forward transfer and resists catastrophic forgetting…

q-bio.QM2022★ 2 cited

Learning multi-scale functional representations of proteins from single-cell microscopy data

Anastasia Razdaibiedina, Alexander Brechalov

Protein function is inherently linked to its localization within the cell, and fluorescent microscopy data is an indispensable resource for learning representations of proteins. De…