2 papers
cs.CV2025
Can Large Vision-Language Models Correct Semantic Grounding Errors By Themselves?
Yuan-Hong Liao, Rafid Mahmood, Sanja Fidler +1
Enhancing semantic grounding abilities in Vision-Language Models (VLMs) often involves collecting domain-specific training data, refining the network architectures, or modifying th…
cs.CV2024
Uncertainty Estimation for 3D Object Detection via Evidential Learning
Nikita Durasov, Rafid Mahmood, Jiwoong Choi +4
3D object detection is an essential task for computer vision applications in autonomous vehicles and robotics. However, models often struggle to quantify detection reliability, lea…