6 papers
FOLIO: Focused Semantic Memory for Streaming Video Understanding
Haoyang Fan, Dhruv Parikh, Anvitha Ramachandran +4
In online streaming video understanding, a video stream continues to arrive and queries may be issued at any time. Because streaming frames grow without bound, the system must cont…
Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics
Anvitha Ramachandran, Dhruv Parikh, Haoyang Fan +2
State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal sequence modeling. While effect…
Can Graphs Help Vision SSMs See Better?
Dhruv Parikh, Anvitha Ramachandran, Haoyang Fan +3
Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends critically on the representation…
SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
Gangda Deng, Zhaoling Chen, Zhongming Yu +11
Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-running systems, are now entrusted to drive…
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
Dhruv Parikh, Haoyang Fan, Rajgopal Kannan +1
Vision-Language Models (VLMs) are expensive because the LLM processes hundreds of largely redundant visual tokens. Existing token reduction methods typically exploit \textit{either…
Tempus Core: Area-Power Efficient Temporal-Unary Convolution Core for Low-Precision Edge DLAs
Prabhu Vellaisamy, Harideep Nair, Thomas Kang +6
The increasing complexity of deep neural networks (DNNs) poses significant challenges for edge inference deployment due to resource and power constraints of edge devices. Recent wo…