8 papers
DamageScope: Vision-Language Retrieval at Scale for Disaster Damage Assessment from Satellite Imagery
Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas +1
Timely and accurate assessment of property damage is critical following natural disasters. Traditional on-site inspections are labor-intensive, costly, and often pose safety risks.…
Open-SAT: LLM-Guided Query Embedding Refinement for Open-Vocabulary Object Retrieval in Satellite Imagery
Md Adnan Arefeen, Biplob Debnath, Ravi K. Rajendran +2
In satellite applications, user queries often take the form of open-ended natural language, extending beyond a fixed set of predefined categories. This open-vocabulary nature poses…
RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution
Arunabh Srivastava, Mohammad A., Khojastepour +2
Humans solve problems by executing targeted plans, yet large language models (LLMs) remain unreliable for structured workflow execution. We propose RunAgent, a multi-agent plan exe…
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
Sarosij Bose, Ravi K. Rajendran, Biplob Debnath +3
Radiology Report Generation (RRG) is a critical step toward automating healthcare workflows, facilitating accurate patient assessments, and reducing the workload of medical profess…
TrafficLens: Multi-Camera Traffic Video Analysis Using LLMs
Md Adnan Arefeen, Biplob Debnath, Srimat Chakradhar
Traffic cameras are essential in urban areas, playing a crucial role in intelligent transportation systems. Multiple cameras at intersections enhance law enforcement capabilities,…
StreamingRAG: Real-time Contextual Retrieval and Generation Framework
Murugan Sankaradas, Ravi K. Rajendran, Srimat T. Chakradhar
Extracting real-time insights from multi-modal data streams from various domains such as healthcare, intelligent transportation, and satellite remote sensing remains a challenge. H…