Publications (54)
Chunkwise Aligners for Streaming Speech Recognition
Wen Shen Teo, Takafumi Moriya, Masato Mimura
We propose the Chunkwise Aligner, a novel architecture for streaming automatic speech recognition (ASR). While the Transducer is the standard model for streaming ASR, its training…
Strong algebraization of fixed point properties
Masato Mimura
The following natural question arises from Shalom's innovational work (1999, Publ. IHES): "Can we establish an intrinsic criterion to synthesize relative fixed point properties int…
Enhancing Monotonic Multihead Attention for Streaming ASR
Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara
We investigate a monotonic multihead attention (MMA) by extending hard monotonic attention to Transformer-based automatic speech recognition (ASR) for online streaming applications…
Avoiding a shape, and the slice rank method for a system of equations
Masato Mimura, Norihide Tokushige
Fix a vector space over a finite field and a system of linear equations. We provide estimates, in terms of the dimension of the vector space, of the maximum of the sizes of subsets…
Group approximation in Cayley topology and coarse geometry, Part I: Coarse embeddings of amenable groups
Masato Mimura, Hiroki Sako
The objective of this series is to study metric geometric properties of (coarse) disjoint unions of amenable Cayley graphs. We employ the Cayley topology and observe connections be…
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
Takafumi Moriya, Masato Mimura, Tomohiro Tanaka +3
This paper proposes a unified framework, All-in-One ASR, that allows a single model to support multiple automatic speech recognition (ASR) paradigms, including connectionist tempor…
Coarse group theoretic study on stable mixed commutator length
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
Let be a group and a normal subgroup of . We study the large scale behavior, not the exact values themselves, of the stable mixed commutator length on the mi…
Constellations in prime elements of number fields
Wataru Kai, Masato Mimura, Akihiro Munemasa +2
Given any number field, we prove that there exist arbitrarily shaped constellations consisting of pairwise non-associate prime elements of the ring of integers. This result extends…
Progressive Alignment Objectives for Aligner-Encoder based ASR
Jaeyoung Lee, Masato Mimura, Takafumi Moriya
Aligner-Encoders are recently proposed seq2seq end-to-end ASR models that replace decoder attention by predicting the uth token directly from the u-th encoder position, so the enco…
Bavard's duality theorem for mixed commutator length
Morimichi Kawasaki, Mitsuaki Kimura, Takahiro Matsushita +1
Let be a normal subgroup of a group . A quasimorphism on is -invariant if for every and every . The goal in this paper is…
Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition
Kazuki Shimada, Yoshiaki Bando, Masato Mimura +3
This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response…
Non-extendablity of Shelukhin's quasimorphism and non-triviality of Reznikov's class
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
Shelukhin constructed a quasimorphism on the universal covering of the group of Hamiltonian diffeomorphisms for a general closed symplectic manifold. In the present paper, we prove…
Group approximation in Cayley topology and coarse geometry, Part III: Geometric property (T)
Masato Mimura, Narutaka Ozawa, Hiroki Sako +1
In this series of papers, we study correspondence between the following: (1) large scale structure of the metric space bigsqcup_m {Cay(G(m))} consisting of Cayley graphs of finite…
Non-autoregressive Error Correction for CTC-based ASR with Phone-conditioned Masked LM
Hayato Futami, Hirofumi Inaguma, Sei Ueno +3
Connectionist temporal classification (CTC) -based models are attractive in automatic speech recognition (ASR) because of their non-autoregressive nature. To take advantage of text…
Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization
Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama +2
This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach…
Commuting symplectomorphisms on a surface and the flux homomorphism
Morimichi Kawasaki, Mitsuaki Kimura, Takahiro Matsushita +1
Let be a closed connected oriented surface whose genus is at least two equipped with a symplectic form. Then we show the vanishing of the cup product of the fluxes of…
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15
We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-en…
Amenability versus non-exactness of dense subgroups of a compact group
Masato Mimura
Given a countable residually finite group, we construct a compact group K and two elements w and u of K with the following properties: The group generated by w and the cube of u is…
Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR
Hirofumi Inaguma, Masato Mimura, Shinsuke Sakai +1
Acoustic-to-word (A2W) end-to-end automatic speech recognition (ASR) systems have attracted attention because of an extremely simplified architecture and fast decoding. To alleviat…
Time-domain Speech Enhancement Assisted by Multi-resolution Frequency Encoder and Decoder
Hao Shi, Masato Mimura, Longbiao Wang +2
Time-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS introduces multi-resolution STFT loss to enhance performance. However, so…
Avoiding a star of three-term arthmetic progressions
Masato Mimura, Norihide Tokushige
We provide an upper bound of the size of a subset A of F_p^n that does not admit a k-star of 3-APs (three-term arithmetic progressions). Namely, the subset A is assumed to contain…
An alternative proof of Kazhdan property for elementary groups
Masato Mimura
In 2010, Invent. Math., Ershov and Jaikin-Zapirain proved Kazhdan's property (T) for elementary groups. This expository article focuses on presenting an alternative simpler proof o…
Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
Kohei Matsuura, Takanori Ashihara, Takafumi Moriya +4
This paper introduces a novel approach called sentence-wise speech summarization (Sen-SSum), which generates text summaries from a spoken document in a sentence-by-sentence manner.…
Invariant quasimorphisms for groups acting on the circle and non-equivalence of SCL
Shuhei Maruyama, Takahiro Matsushita, Masato Mimura
We construct invariant quasimorphisms for groups acting on the circle. Furthermore, we provide a criterion for the non-extendablity of the resulting quasimorphisms and an explicit…
Survey on invariant quasimorphisms and stable mixed commutator length
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
A homogeneous quasimorphism on a normal subgroup of is said to be -invariant if for every and for every . Invariant quasim…
Mixed commutator lengths, wreath products and general ranks
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
In the present paper, for a pair of a group and its normal subgroup , we consider the mixed commutator length on the mixed commutator subgroup $[…
On strong property (T) and fixed point properties for Lie groups
Tim de Laat, Masato Mimura, Mikael de la Salle
We consider certain strengthenings of property (T) relative to Banach spaces that are satisfied by high rank Lie groups. Let X be a Banach space for which, for all k, the Banach--M…
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando +15
In this paper, we introduce a multi-talker distant automatic speech recognition (DASR) system we designed for the DASR task 1 of the CHiME-8 challenge. Our system performs speaker…
Group approximation in Cayley topology and coarse geometry, Part II: Fibered coarse embeddings
Masato Mimura, Hiroki Sako
The objective of this series is to study metric geometric properties of disjoint unions of amenable Cayley graphs by group properties of the Cayley accumulation points in the space…
Alignment-Free Training for Transducer-based Multi-Talker ASR
Takafumi Moriya, Shota Horiguchi, Marc Delcroix +5
Extending the RNN Transducer (RNNT) to recognize multi-talker speech is essential for wider automatic speech recognition (ASR) applications. Multi-talker RNNT (MT-RNNT) aims to ach…
CTC-synchronous Training for Monotonic Attention Model
Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara
Monotonic chunkwise attention (MoChA) has been studied for the online streaming automatic speech recognition (ASR) based on a sequence-to-sequence framework. In contrast to connect…
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
Jaeyoung Lee, Masato Mimura
We present a decoder-only Conformer for automatic speech recognition (ASR) that processes speech and text in a single stack without external speech encoders or pretrained large lan…
Flux homomorphism and bilinear form constructed from Shelukhin's quasimorphism
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
Given a closed connected symplectic manifold , we construct an alternating -bilinear form on the real first cohom…
End-to-end Music-mixed Speech Recognition
Jeongwoo Woo, Masato Mimura, Kazuyoshi Yoshii +1
Automatic speech recognition (ASR) in multimedia content is one of the promising applications, but speech data in this kind of content are frequently mixed with background music, w…
Distilling the Knowledge of BERT for Sequence-to-Sequence ASR
Hayato Futami, Hirofumi Inaguma, Sei Ueno +3
Attention-based sequence-to-sequence (seq2seq) models have achieved promising results in automatic speech recognition (ASR). However, as these models decode in a left-to-right way,…
Multi-way expanders and imprimitive group actions on graphs
Masato Mimura
For n at least 2, the concept of n-way expanders was defined by various researchers. Bigger n gives a weaker notion in general, and 2-way expanders coincide with expanders in usual…
An extreme counterexample to the Lubotzky--Weiss conjecture
Masato Mimura
In 1993, Lubotzky and Weiss conjectured that if a compact group admits two finitely generated dense subgroups, one of which is amenable and the other has Kazhdan's property (T), th…
Sphere equivalence, Banach expanders, and extrapolation
Masato Mimura
We study the Banach spectral gap lambda_1(G;X,p) of finite graphs G for pairs (X,p) of Banach spaces and exponents. We define the notion of sphere equivalence between Banach spaces…
ASR Rescoring and Confidence Estimation with ELECTRA
Hayato Futami, Hirofumi Inaguma, Masato Mimura +2
In automatic speech recognition (ASR) rescoring, the hypothesis with the fewest errors should be selected from the n-best list using a language model (LM). However, LMs are usually…
Generative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech Recognition
Kohei Matsuura, Masato Mimura, Shinsuke Sakai +1
It is important to transcribe and archive speech data of endangered languages for preserving heritages of verbal culture and automatic speech recognition (ASR) is a powerful tool t…
Superrigidity from Chevalley groups into acylindrically hyperbolic groups via quasi-cocycles
Masato Mimura
We prove that every homomorphism from the elementary Chevalley group over a finitely generated unital commutative ring associated with reduced irreducible classical root system of…
Fixed point properties and second bounded cohomology of universal lattices on Banach space
Masato Mimura
Let B be any Lp space for p in (1,infty) or any Banach space isomorphic to a Hilbert space, and k be a nonnegative integer. We show that if n is at least 4, then the universal latt…
Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu Language
Kohei Matsuura, Sei Ueno, Masato Mimura +2
Ainu is an unwritten language that has been spoken by Ainu people who are one of the ethnic groups in Japan. It is recognized as critically endangered by UNESCO and archiving and d…
Improving Large-Scale Weakly Supervised ASR by Filtering and Selection
Kohei Matsuura, Masato Mimura
Leveraging large-scale weakly supervised datasets is crucial to train robust end-to-end automatic speech recognition (ASR) models. However, such datasets often contain noisy labels…
Fixed point property for universal lattice on Schatten classes
Masato Mimura
The special linear group G=SL_n(Z[x1,...,xk]) (n at least 3 and k finite) is called the universal lattice. Let n be at least 4, p be any real number in (1,\infty). The main result…
Invariant quasimorphisms and generalized mixed Bavard duality
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
This article provides an expository account of the celebrated duality theorem of Bavard and three its strengthenings. The Bavard duality theorem connects scl (stable commutator len…
The space of non-extendable quasimorphisms
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
For a pair of a group and its normal subgroup , we consider the space of quasimorphisms and quasi-cocycles on non-extendable to . To treat this space, we esta…
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
Takafumi Moriya, Takanori Ashihara, Masato Mimura +4
A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel int…
Distilling the Knowledge of BERT for CTC-based ASR
Hayato Futami, Hirofumi Inaguma, Masato Mimura +2
Connectionist temporal classification (CTC) -based models are attractive because of their fast inference in automatic speech recognition (ASR). Language model (LM) integration appr…
On Quasi-homomorphisms and Commutators in the Special Linear Group over a Euclidean Ring
Masato Mimura
We prove that for any euclidean ring R and n at least 6, Gamma=SL_n(R) has no unbounded quasi-homomorphisms. From Bavard's duality theorem, this means that the stable commutator le…
Coarse geometry of stable mixed commutator length I: duality and functional analysis on chains
Morimichi Kawasaki, Mitsuaki Kimura, Shuhei Maruyama +2
Let be a group and its normal subgroup. On the mixed commutator subgroup , the mixed stable commutator length and the restriction of the ordinar…
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
Hiroshi Sato, Takafumi Moriya, Masato Mimura +6
Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real…
Property modulo and homomorphism superrigidity into mapping class groups
Masato Mimura
Every homomorphism from finite index subgroups of a universal lattices to mapping class groups of orientable surfaces (possibly with punctures), or to outer automorphism groups of…
On the spectrum and linear programming bound for hypergraphs
Sebastian M. CioabÄ, Jack H. Koolen, Masato Mimura +2
The spectrum of a graph is closely related to many graph parameters. In particular, the spectral gap of a regular graph which is the difference between its valency and second eigen…