6 papers · 1 filter
CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning
Andrew Kiruluta, Preethi Raju, Priscilla Burity
Vision-Language Models (vLLMs) have emerged as powerful architectures for joint reasoning over visual and textual inputs, enabling breakthroughs in image captioning, cross modal re…
From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs
Andrew Kiruluta, Priscilla Burity
Vision-language models (VLMs) unify computer vision and natural language processing in a single architecture capable of interpreting and describing images. Most state-of-the-art sy…
Spectral Dictionary Learning for Generative Image Modeling
Andrew Kiruluta
We propose a novel spectral generative model for image synthesis that departs radically from the common variational, adversarial, and diffusion paradigms. In our approach, images,…
Wavelet-based Variational Autoencoders for High-Resolution Image Generation
Andrew Kiruluta
Variational Autoencoders (VAEs) are powerful generative models capable of learning compact latent representations. However, conventional VAEs often generate relatively blurry image…
A Hybrid Wavelet-Fourier Method for Next-Generation Conditional Diffusion Models
Andrew Kiruluta, Andreas Lemos
We present a novel generative modeling framework,Wavelet-Fourier-Diffusion, which adapts the diffusion paradigm to hybrid frequency representations in order to synthesize high-qual…
Hierarchical Attention Diffusion Networks with Object Priors for Video Change Detection
Andrew Kiruluta, Eric Lundy, Andreas Lemos
We present a unified change detection pipeline that combines instance level masking, multi\-scale attention within a denoising diffusion model, and per pixel semantic classificatio…