2 papers
cs.LG2026
FAAR: Format-Aware Adaptive Rounding for NVFP4
Hanglin Li, Shuchang Tian, Chen Lin +2
Deploying large language models (LLMs) on edge devices requires extremely low-bit quantization. Ultra-low precision formats such as NVFP4 offer a promising solution for reducing me…
cs.CV2026
Efficient Token Pruning for LLaDA-V
Zhewen Wan, Tianchen Song, Chen Lin +2
Diffusion-based large multimodal models, such as LLaDA-V, have demonstrated impressive capabilities in vision-language understanding and generation. However, their bidirectional at…