14 papers
Can Graphs Help Vision SSMs See Better?
Dhruv Parikh, Anvitha Ramachandran, Haoyang Fan +3
Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends critically on the representation…
Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery
Dennis Menn, Yuedong Yang, Bokun Wang +6
Current video generation models suffer from high computational latency, making real-time applications prohibitively costly. In this paper, we address this limitation by exploiting…
Parameter Efficient Fine-tuning for Domain-specific Gastrointestinal Disease Recognition
Sanjaya Poudel, Nikita Kunwor, Raj Simkhada +3
Despite recent advancements in the field of medical image analysis with the use of pretrained foundation models, the issue of distribution shifts between cross-source images largel…
Single-Round Scalable Analytic Federated Learning
Alan T. L. Bacellar, Mustafa Munir, Felipe M. G. França +3
Federated Learning (FL) is plagued by two key challenges: high communication overhead and performance collapse on heterogeneous (non-IID) data. Analytic FL (AFL) provides a single-…
Fuel Gauge: Estimating Chain-of-Thought Length Ahead of Time in Large Multimodal Models
Yuedong Yang, Xiwen Wei, Mustafa Munir +1
Reasoning Large Multi-modality Models (LMMs) have become the de facto choice for many applications. However, these models rely on a Chain-of-Thought (CoT) process that is lengthy a…
PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
Mustafa Munir, Md Mostafijur Rahman, Kartikeya Bhardwaj +2
Long-form video editing poses unique challenges due to the exponential increase in the computational cost from joint editing and Denoising Diffusion Implicit Models (DDIM) inversio…