5 papers
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans
Nabaraj Subedi, Shuvo Dip Datta, Ahmed Abdelaty +1
Civil infrastructure compliance checking has long relied on engineers manually reading legacy 2D plans; however, OCR-based automation strips away the geometry and layout essential…
Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges
Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar
Object-counting methods have rapidly shifted from class-specific density regression to open-vocabulary, foundation-model-backed counters. These methods now enumerate instances from…
Parameter Golf: What Really Works?
Prashanna Mani Paudel, Shivanand Venkanna Sheshappanavar
How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which participants trained the best langua…
When More Documents Hurt RAG: Mitigating Vector Search Dilution with Domain-Scoped, Model-Agnostic Retrieval
Nabaraj Subedi, Ahmed Abdelaty, Shivanand Venkanna Sheshappanavar
Retrieval-augmented generation degrades when scaled to large, heterogeneous document collections, where dense similarity loses discriminative power, and top-k retrieval increasingl…
The Effect of Negation on CLIP in Medical Imaging: Limitations of Contrastive Language-Image Pretraining
Jasmine Vu, Shivanand Sheshappanavar
Large vision-language models like CLIP are increasingly used in medical imaging tasks due to their ability to align images and text without the need for extensive labeled data. Thi…