14 citations · 16 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
Haojun Xia, Xiaoxia Wu, Jisen Li +12
The KV cache is a dominant memory bottleneck for LLM inference. While 4-bit KV quantization preserves accuracy, 2-bit often degrades it, especially on long-context reasoning. We cl…
cs.LG2020★ 14 cited
Predicting Future Sales of Retail Products using Machine Learning
Devendra Swami, Alay Dilipbhai Shah, Subhrajeet K B Ray
Techniques for making future predictions based upon the present and past data, has always been an area with direct application to various real life problems. We are discussing a si…