2 citations · 3 across the 6 of their papers we have counts for
3 papers · 1 filter
Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?
Fumihiko Tachibana, Daisuke Miyashita, Jun Deguchi
In Retrieval-Augmented Generation (RAG) systems, a large number of retrieved chunks are concatenated to form the input context so that users can receive high-quality responses base…
Residual-Mass Accounting for Partial-KV Decoding
Yasuto Hoshi, Daisuke Miyashita, Jun Deguchi
We study a controlled partial-KV decoding setting in which exact unnormalized softmax contributions are computed for sink/tail anchors and a retrieved token set, while the remainin…
On Storage Neural Network Augmented Approximate Nearest Neighbor Search
Taiga Ikeda, Daisuke Miyashita, Jun Deguchi
Large-scale approximate nearest neighbor search (ANN) has been gaining attention along with the latest machine learning researches employing ANNs. If the data is too large to fit i…