1 paper
Yusheng Tian, Jingyu Li, Tan Lee
Pooling is needed to aggregate frame-level features into utterance-level representations for speaker modeling. Given the success of statistics-based pooling methods, we hypothesize…