5 papers
SoftMatcha 2: A Fast and Soft Pattern Matcher for Trillion-Scale Corpora
Masataka Yoneda, Yusuke Matsushita, Go Kamoda +4
We present SoftMatcha 2, an ultra-fast and flexible search algorithm that enables search over trillion-scale natural language corpora in under 0.3 seconds while allowing semantic v…
Quantifying Lexical Semantic Shift via Unbalanced Optimal Transport
Ryo Kishino, Hiroaki Yamagiwa, Ryo Nagata +2
Lexical semantic change detection aims to identify shifts in word meanings over time. While existing methods using embeddings from a diachronic corpus pair estimate the degree of c…
SoftMatcha: A Soft and Fast Pattern Matcher for Billion-Scale Corpus Searches
Hiroyuki Deguchi, Go Kamoda, Yusuke Matsushita +4
Researchers and practitioners in natural language processing and computational linguistics frequently observe and analyze the real language usage in large-scale corpora. For that p…
TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
Makoto Shing, Kou Misaki, Han Bao +2
Causal language models have demonstrated remarkable capabilities, but their size poses significant challenges for deployment in resource-constrained environments. Knowledge distill…
Zipfian Whitening
Sho Yokoi, Han Bao, Hiroto Kurita +1
The word embedding space in neural models is skewed, and correcting this can improve task performance. We point out that most approaches for modeling, correcting, and measuring the…