3 papers
cs.LG2026
Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference
Yifei Gao, Lei Wang, Rong-Cheng Tu +3
A core bottleneck in large language model (LLM) inference is the cost of attending over the ever-growing key-value (KV) cache. Although near-oracle top-k KV selection can preserve…
cs.CL2025
Compensate Quantization Errors+: Quantized Models Are Inquisitive Learners
Yifei Gao, Jie Ou, Lei Wang +2
The quantization of large language models (LLMs) has been a prominent research area aimed at enabling their lightweight deployment in practice. Existing research about LLM's quanti…
cs.GR2025
Beyond Existance: Fulfill 3D Reconstructed Scenes with Pseudo Details
Yifei Gao, Jun Huang, Lei Wang +2
The emergence of 3D Gaussian Splatting (3D-GS) has significantly advanced 3D reconstruction by providing high fidelity and fast training speeds across various scenarios. While rece…