7 papers · 1 filter
CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning
Andrew Kiruluta, Preethi Raju, Priscilla Burity
Vision-Language Models (vLLMs) have emerged as powerful architectures for joint reasoning over visual and textual inputs, enabling breakthroughs in image captioning, cross modal re…
From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs
Andrew Kiruluta, Priscilla Burity
Vision-language models (VLMs) unify computer vision and natural language processing in a single architecture capable of interpreting and describing images. Most state-of-the-art sy…
Hierarchical Attention Diffusion Networks with Object Priors for Video Change Detection
Andrew Kiruluta, Eric Lundy, Andreas Lemos
We present a unified change detection pipeline that combines instance level masking, multi\-scale attention within a denoising diffusion model, and per pixel semantic classificatio…
Spectral Dictionary Learning for Generative Image Modeling
Andrew Kiruluta
We propose a novel spectral generative model for image synthesis that departs radically from the common variational, adversarial, and diffusion paradigms. In our approach, images,…
Reducing Deep Network Complexity via Sparse Hierarchical Fourier Interaction Networks
Andrew Kiruluta, Samantha Williams
This paper presents a Sparse Hierarchical Fourier Interaction Networks, an architectural building block that unifies three complementary principles of frequency domain modeling: A…
Wavelet-based Variational Autoencoders for High-Resolution Image Generation
Andrew Kiruluta
Variational Autoencoders (VAEs) are powerful generative models capable of learning compact latent representations. However, conventional VAEs often generate relatively blurry image…