4 papers
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
Yifei Gao, Lei Wang, Rong-Cheng Tu +3
A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve…
Beyond Existance: Fulfill 3D Reconstructed Scenes with Pseudo Details
Yifei Gao, Jun Huang, Lei Wang +2
The emergence of 3D Gaussian Splatting (3D-GS) has significantly advanced 3D reconstruction by providing high fidelity and fast training speeds across various scenarios. While rece…
Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners
Yifei Gao, Jie Ou, Lei Wang +2
The quantization of large language models (LLMs) has been a prominent research area aimed at enabling their lightweight deployment in practice. Existing research about LLM's quanti…
Compensate Quantization Errors: Make Weights Hierarchical to Compensate Each Other
Yifei Gao, Jie Ou, Lei Wang +4
Emergent Large Language Models (LLMs) use their extraordinary performance and powerful deduction capacity to discern from traditional language models. However, the expenses of comp…