14 citations · 39 across the 14 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.DC2026
FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim +4
Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual grounding, offering a natural perception interf…
cs.CV2026
VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
Rahul Atul Bhope, K. R. Jayaram, Vinod Muthusamy +3
Despite significant costs from retrieving and processing high-fidelity visual inputs, most multimodal vision-language systems operate at fixed fidelity levels. We introduce VOILA,…