Publications (6)
Shifts: A Dataset of Real Distributional Shift Across Multiple Large-Scale Tasks
Andrey Malinin, Neil Band, Ganshin +15
There has been significant research done on developing methods for improving robustness to distributional shift and uncertainty estimation. In contrast, only limited work has exami…
Multi-Sentence Resampling: A Simple Approach to Alleviate Dataset Length Bias and Beam-Search Degradation
Ivan Provilkov, Andrey Malinin
Neural Machine Translation (NMT) is known to suffer from a beam-search problem: after a certain point, increasing beam size causes an overall drop in translation quality. This effe…
Vertex and Energy Reconstruction in JUNO with Machine Learning Methods
Zhen Qian, Vladislav Belavin, Vasily Bokov +20
The Jiangmen Underground Neutrino Observatory (JUNO) is an experiment designed to study neutrino oscillations. Determination of neutrino mass ordering and precise measurement of ne…
BPE-Dropout: Simple and Effective Subword Regularization
Ivan Provilkov, Dmitrii Emelianenko, Elena Voita
Subword segmentation is widely used to address the open vocabulary problem in machine translation. The dominant approach to subword segmentation is Byte Pair Encoding (BPE), which…
Escaping the Verifier: Learning to Reason via Demonstrations
Locke Cai, Max Ryabinin, Ivan Provilkov
Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-world reasoning-intensive tasks lack ve…
Regression Prior Networks
Andrey Malinin, Sergey Chervontsev, Ivan Provilkov +1
Prior Networks are a recently developed class of models which yield interpretable measures of uncertainty and have been shown to outperform state-of-the-art ensemble approaches on…