3 papers
cs.IR2026
The Pre-Training Study of Expanded-SPLADE Models on Web Document Titles
Hiun Kim, Tae Kwan Lee, Taeryun Won
Masked Language Modeling (MLM) pre-training is one of the primary ways to initialize Neural Information Retrieval (IR) models prior to retrieval fine-tuning. However, studies show…
cs.IR2026
The Role of Vocabularies in Learning Sparse Representations for Ranking
Hiun Kim, Tae Kwan Lee, Taeryun Won
Learned Sparse Retrieval (LSR) such as SPLADE has growing interest for effective semantic 1st stage matching while enjoying the efficiency of inverted indices. A recent work on lea…
cs.IR2025
Efficiency and Effectiveness of SPLADE Models on Billion-Scale Web Document Title
Taeryun Won, Tae Kwan Lee, Hiun Kim +1
This paper presents a comprehensive comparison of BM25, SPLADE, and Expanded-SPLADE models in the context of large-scale web document retrieval. We evaluate the effectiveness and e…