3 papers
cs.CV2025
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
Wufei Ma, Luoxin Ye, Celso M de Melo +2
Humans naturally understand 3D spatial relationships, enabling complex reasoning like predicting collisions of vehicles from different directions. Current large multimodal models (…
cs.CV2025
SHARDeg: A Benchmark for Skeletal Human Action Recognition in Degraded Scenarios
Simon Malzard, Nitish Mital, Richard Walters +3
Computer vision (CV) models for detection, prediction or classification tasks operate on video data-streams that are often degraded in the real world, due to deployment in real-tim…
cs.CV2025
Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data
Prasanna Reddy Pulakurthi, Majid Rabbani, Celso M. de Melo +2
This paper introduces a novel dual-region augmentation approach designed to reduce reliance on large-scale labeled datasets while improving model robustness and adaptability across…