2 papers
cs.CV2025
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
Zachary Huemann, Samuel Church, Joshua D. Warner +7
Vision-language models can connect the text description of an object to its specific location in an image through visual grounding. This has potential applications in enhanced radi…
cs.RO2024
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
Matthew Chang, Gunjan Chhablani, Alexander Clegg +17
We present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhib…