14 papers
When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression
Riku Kisako, Hayato Tsukagoshi, Ryohei Sasano
Recent high-performing text embedding models often output high-dimensional real-valued vectors, resulting in substantial storage and computational costs. To address this issue, com…
Can We Still Hear the Accent? Investigating the Resilience of Native Language Signals in the LLM Era
Nabelanita Utami, Ryohei Sasano
The evolution of writing assistance tools from machine translation to large language models (LLMs) has changed how researchers write. This study investigates whether this shift is…
How Do Language Models Acquire Character-Level Information?
Soma Sato, Ryohei Sasano
Language models (LMs) have been reported to implicitly encode character-level information, despite not being explicitly provided during training. However, the mechanisms underlying…
Do LLMs and Humans Find the Same Questions Difficult? A Case Study on Japanese Quiz Answering
Naoya Sugiura, Kosuke Yamada, Yasuhiro Ogawa +2
LLMs have achieved performance that surpasses humans in many NLP tasks. However, it remains unclear whether problems that are difficult for humans are also difficult for LLMs. This…
FrameEOL: Semantic Frame Induction using Causal Language Models
Chihiro Yano, Kosuke Yamada, Hayato Tsukagoshi +2
Semantic frame induction is the task of clustering frame-evoking words according to the semantic frames they evoke. In recent years, leveraging embeddings of frame-evoking words th…
Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings
Hayato Tsukagoshi, Ryohei Sasano
Prompt-based text embedding models, which generate task-specific embeddings upon receiving tailored prompts, have recently demonstrated remarkable performance. However, their resul…