4 citations · 9 across the 9 of their papers we have counts for
6 papers · 1 filter
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
Janak Kapuriya, Anwar Shaikh, Arnav Goel +8
In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLM…
Generative AI in Vision: A Survey on Models, Metrics and Applications
Gaurav Raut, Apoorv Singh
Generative AI models have revolutionized various fields by enabling the creation of realistic and diverse data samples. Among these models, diffusion models have emerged as a power…
Training Strategies for Vision Transformers for Object Detection
Apoorv Singh
Vision-based Transformer have shown huge application in the perception module of autonomous driving in terms of predicting accurate 3D bounding boxes, owing to their strong capabil…
Surround-View Vision-based 3D Detection for Autonomous Driving: A Survey
Apoorv Singh, Varun Bankiti
Vision-based 3D Detection task is fundamental task for the perception of an autonomous driving system, which has peaked interest amongst many researchers and autonomous driving eng…
Transformer-Based Sensor Fusion for Autonomous Driving: A Survey
Apoorv Singh
Sensor fusion is an essential topic in many perception systems, such as autonomous driving and robotics. Transformers-based detection head and CNN-based feature encoder to extract…
Vision-RADAR fusion for Robotics BEV Detections: A Survey
Apoorv Singh
Due to the trending need of building autonomous robotic perception system, sensor fusion has attracted a lot of attention amongst researchers and engineers to make best use of cros…