12 papers
LTM: Large-scale Terrain Model for Wildfire-prone Landscapes
Xiao Fu, Yue Hu, Meida Chen +2
Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards. However, wildfire-prone regions often span vast areas where conventional reconstructi…
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
Youngrae Kim, Qixin Hu, C. -C. Jay Kuo +1
Autoregressive diffusion enables real-time frame streaming, yet existing sliding-window caches discard past context, causing fidelity degradation, identity drift, and motion stagna…
RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
Jingqi Xu, Jingxi Lu, Chenghao Li +3
Vision-Language Models (VLMs) have achieved remarkable progress in multimodal reasoning and generation, yet their high computational demands remain a major challenge. Diffusion Vis…
LAWCAT: Efficient Distillation from Quadratic to Linear Attention with Convolution across Tokens for Long Context Modeling
Zeyu Liu, Souvik Kundu, Lianghao Jiang +5
Although transformer architectures have achieved state-of-the-art performance across diverse domains, their quadratic computational complexity with respect to sequence length remai…
Bitwidth-Specific Logarithmic Arithmetic for Future Hardware-Accelerated Training
Hassan Hamad, Yuou Qiu, Peter A. Beerel +1
While advancements in quantization have significantly reduced the computational costs of inference in deep learning, training still predominantly relies on complex floating-point a…
Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
Sreetama Sarkar, Yue Che, Alex Gavin +2
Despite their remarkable progress in multimodal understanding tasks, large vision language models (LVLMs) often suffer from "hallucinations", generating texts misaligned with the v…