9 papers
PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision Tasks
Deepank Singh, Anurag Nihal, Vedhus Hoskere
Self-supervised Vision Foundation Models (VFMs) have become essential backbones for downstream tasks due to their strong and transferable visual representations. However, their pat…
BridgeEQA: Virtual Embodied Agents for Real Bridge Inspections
Subin Varghese, Joshua Gao, Asad Ur Rahman +1
Deploying embodied agents that can answer questions about their surroundings in realistic real-world settings remains difficult, partly due to the scarcity of benchmarks for episod…
Excite, Attend and Segment (EASe): Domain-Agnostic Fine-Grained Mask Discovery with Feature Calibration and Self-Supervised Upsampling
Deepank Singh, Anurag Nihal, Vedhus Hoskere
Unsupervised segmentation approaches have increasingly leveraged foundation models (FM) to improve salient object discovery. However, these methods often falter in scenes with comp…
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
Joshua Gao, Quoc Huy Pham, Subin Varghese +2
Retrieval-Augmented Generation (RAG) is a critical technique for grounding Large Language Models (LLMs) in factual evidence, yet evaluating RAG systems in specialized, safety-criti…
ViewDelta: Scaling Scene Change Detection through Text-Conditioning
Subin Varghese, Joshua Gao, Vedhus Hoskere
We introduce a generalized framework for Scene Change Detection (SCD) that addresses the core ambiguity of distinguishing "relevant" from "nuisance" changes, enabling effective joi…
Vision-Based Adaptive Robotics for Autonomous Surface Crack Repair
Joshua Genova, Eric Cabrera, Vedhus Hoskere
Surface cracks in infrastructure can lead to severe deterioration and expensive maintenance if not efficiently repaired. Manual repair methods are labor-intensive, time-consuming,…