activity
20232026
most citedReshape Dimensions Network for Speaker Recognition

24 citations · 25 across the 5 of their papers we have counts for

collaborators

6 papers

eess.AS2026

VoXtream2: Full-stream TTS with dynamic speaking rate control

Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze

Full-stream text-to-speech (TTS) for interactive systems must start speaking with minimal delay while remaining controllable as text arrives incrementally. We present VoXtream2, a…

eess.AS2025

VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency

Nikita Torgashov, Gustav Eje Henter, Gabriel Skantze

We present VoXtream, a fully autoregressive, zero-shot streaming text-to-speech (TTS) system for real-time use that begins speaking from the first word. VoXtream directly maps inco…

eess.AS2024

Study on Inter and Intra Speaker Variability in Speaker Recognition

Anton Okhotnikov, Nikita Torgashov, Ivan Yakovlev +2

Optimization of a trade-off between the number of speakers and their temporal variability (or session diversity) is crucial for the development of a speaker recognition system toge…

eess.AS202424 cited

Reshape Dimensions Network for Speaker Recognition

Ivan Yakovlev, Rostislav Makarov, Andrei Balykin +3

In this paper, we present Reshape Dimensions Network (ReDimNet), a novel neural network architecture for extracting utterance-level speaker representations. Our approach leverages…

eess.AS20231 cited

LRPD: Large Replay Parallel Dataset

Ivan Yakovlev, Mikhail Melnikov, Nikita Bukhal +4

The latest research in the field of voice anti-spoofing (VAS) shows that deep neural networks (DNN) outperform classic approaches like GMM in the task of presentation attack detect…

eess.AS2023

The ID R&D VoxCeleb Speaker Recognition Challenge 2023 System Description

Nikita Torgashov, Rostislav Makarov, Ivan Yakovlev +3

This report describes ID R&D team submissions for Track 2 (open) to the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). Our solution is based on the fusion of deep ResNets…