2 citations · 2 across the 19 of their papers we have counts for
9 papers · 1 filter
FOLIO: Focused Semantic Memory for Streaming Video Understanding
Haoyang Fan, Dhruv Parikh, Anvitha Ramachandran +4
In online streaming video understanding, a video stream continues to arrive and queries may be issued at any time. Because streaming frames grow without bound, the system must cont…
Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics
Anvitha Ramachandran, Dhruv Parikh, Haoyang Fan +2
State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal sequence modeling. While effect…
Can Graphs Help Vision SSMs See Better?
Dhruv Parikh, Anvitha Ramachandran, Haoyang Fan +3
Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends critically on the representation…
GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA
Anvitha Ramachandran, Dhruv Parikh, Viktor Prasanna
Vision Graph Neural Networks (ViGs) represent an image as a graph of patch tokens, enabling adaptive, feature-driven neighborhoods. Unlike CNNs with fixed grid biases or Vision Tra…
ImageHD: Energy-Efficient On-Device Continual Learning of Visual Representations via Hyperdimensional Computing
Jebacyril Arockiaraj, Dhruv Parikh, Viktor Prasanna
On-device continual learning (CL) is critical for edge AI systems operating on non-stationary data streams, but most existing methods rely on backpropagation or exemplar-heavy clas…
Latent Denoising Improves Visual Alignment in Large Multimodal Models
Dhruv Parikh, Jacob Fein-Ashley, Rajgopal Kannan +1
Large Multimodal Models (LMMs) such as LLaVA are typically trained with an autoregressive language modeling objective, providing only indirect supervision to visual tokens. This of…