18 papers
GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction
Leyang Chen, Junyi Wu, Zhiteng Li +1
Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value (KV) cache that grows linearly with sequence length, creating a severe memory bottl…
FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing
Junyi Wu, Zhiteng Li, Haotong Qin +2
Text-guided image editing with diffusion models has achieved remarkable quality but often suffers from prohibitive latency. We introduce \textbf{FlashEdit}, a real-time localized i…
SOAR: Scale Optimization for Accurate Reconstruction in NVFP4 Quantization
Chengzhu Bao, Xianglong Yan, Zhiteng Li +3
NVFP4 has recently emerged as an efficient 4-bit microscaling format for large language models (LLMs), offering superior numerical fidelity with native hardware support. However, e…
FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching
Yixin Tang, Jiawei Guo, Junxian Li +6
Recently, diffusion-based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, they indiscriminately denoise…
DVD-Quant: Data-free Video Diffusion Transformers Quantization
Zhiteng Li, Hanxuan Li, Junyi Wu +6
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While…
DQuant: Accurate Low-bit Post-Training Weight Quantization for LLMs
Xianglong Yan, ChengZhu Bao, Zhiteng Li +5
Large language models (LLMs) deliver strong performance, but their high compute and memory costs make deployment difficult in resource-constrained scenarios. Weight-only post-train…