activity
20242026
collaborators

7 papers

cs.CV2026

A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving

Jingtao Sun, Xiaohai He, Yike Zhang +4

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a…

cs.CV2025

SDFA: Structure Aware Discriminative Feature Aggregation for Efficient Human Fall Detection in Video

Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian

Older people are susceptible to fall due to instability in posture and deteriorating health. Immediate access to medical support can greatly reduce repercussions. Hence, there is a…

cs.CV2025

Modeling Human Skeleton Joint Dynamics for Fall Detection

Sania Zahan, Ghulam Mubashar Hassan, Ajmal Mian

The increasing pace of population aging calls for better care and support systems. Falling is a frequent and critical problem for elderly people causing serious long-term health is…

cs.CV2025

Multiview Point Cloud Registration Based on Minimum Potential Energy for Free-Form Blade Measurement

Zijie Wu, Yaonan Wang, Yang Mo +5

Point cloud registration is an essential step for free-form blade reconstruction in industrial measurement. Nonetheless, measuring defects of the 3D acquisition system unavoidably…

cs.CV2024

Simultaneous Multiple Object Detection and Pose Estimation using 3D Model Infusion with Monocular Vision

Congliang Li, Shijie Sun, Xiangyu Song +3

Multiple object detection and pose estimation are vital computer vision tasks. The latter relates to the former as a downstream problem in applications such as robotics and autonom…

cs.CL2024

A Comprehensive Overview of Large Language Models

Humza Naveed, Asad Ullah Khan, Shi Qiu +6

Large Language Models (LLMs) have recently demonstrated remarkable capabilities in natural language processing tasks and beyond. This success of LLMs has led to a large influx of r…