10 papers
Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization
Youyuan Liu, Longtao Zhang, Ruoyu Li +6
With the rapid advancement of large-scale scientific simulations, the massive volume of point cloud data generated has increasingly become a critical bottleneck for scientific stor…
3D Gaussian Splatting for Scientific Particle Data Compression and Rendering
Bo Jiang, Youyuan Liu, Taolue Yang +2
Large-scale particle simulations produce hundreds of millions of particles, straining storage, transfer, and interactive visualization. Existing lossy compressors such as SZ3 opera…
FLARE: A Dataflow-Aware and Scalable Hardware Architecture for Neural-Hybrid Scientific Lossy Compression
Wenqi Jia, Zhewen Hu, Baixi Sun +9
The paper introduces FLARE, a hardware architecture that integrates neural network‑based lossy compression with traditional scientific data processing to reduce memory traffic and…
KVSculpt: KV Cache Compression as Distillation
Bo Jiang, Sian Jin
KV cache compression is critical for efficient long-context LLM inference. Approaches that reduce the per-pair footprint -- quantization and low-rank decomposition -- are orthogona…
Event-VStream: Event-Driven Real-Time Understanding for Long Video Streams
Zhenghui Guo, Yuanbin Man, Junyuan Sheng +8
Real-time understanding of long video streams remains challenging for multimodal large language models (VLMs) due to redundant frame processing and rapid forgetting of past context…
PackKV: Reducing KV Cache Memory Footprint through LLM-Aware Lossy Compression
Bo Jiang, Taolue Yang, Youyuan Liu +3
Transformer-based large language models (LLMs) have demonstrated remarkable potential across a wide range of practical applications. However, long-context inference remains a signi…