7 papers
A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving
Jingtao Sun, Xiaohai He, Yike Zhang +4
Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a…
SDFA: Structure Aware Discriminative Feature Aggregation for Efficient Human Fall Detection in Video
Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian
Older people are susceptible to fall due to instability in posture and deteriorating health. Immediate access to medical support can greatly reduce repercussions. Hence, there is a…
Modeling Human Skeleton Joint Dynamics for Fall Detection
Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian
The increasing pace of population aging calls for better care and support systems. Falling is a frequent and critical problem for elderly people causing serious long-term health is…
Multiview Point Cloud Registration Based on Minimum Potential Energy for Free-Form Blade Measurement
Zijie Wu, Yaonan Wang, Yang Mo +5
Point cloud registration is an essential step for free-form blade reconstruction in industrial measurement. Nonetheless, measuring defects of the 3D acquisition system unavoidably…
Simultaneous Multiple Object Detection and Pose Estimation using 3D Model Infusion with Monocular Vision
Congliang Li, Shijie Sun, Xiangyu Song +3
Multiple object detection and pose estimation are vital computer vision tasks. The latter relates to the former as a downstream problem in applications such as robotics and autonom…
A Comprehensive Overview of Large Language Models
Humza Naveed, Asad Ullah Khan, Shi Qiu +6
Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of r…