9 papers
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
Aritra Dutta, Somak Aditya
Vision-language models (VLMs) increasingly rely on chain-of-thought (CoT) reasoning to solve complex multimodal tasks, but their large parameter sizes make deployment expensive. St…
AFGNN: API Misuse Detection using Graph Neural Networks and Clustering
Ponnampalam Pirapuraj, Tamal Mondal, Sharanya Gupta +3
Application Programming Interfaces (APIs) are crucial to software development, enabling integration of existing systems with new applications by reusing tried and tested code, savi…
AD: Analysis and Detection of Adversarial Threats in Visual Perception for End-to-End Autonomous Driving Systems
Ishan Sahu, Somnath Hazra, Somak Aditya +1
End-to-end autonomous driving systems have achieved significant progress, yet their adversarial robustness remains largely underexplored. In this work, we conduct a closed-loop eva…
PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
Sachin Vashistha, Aryan Bibhuti, Atharva Naik +2
Real-world conversations are rich with pragmatic elements, such as entity mentions, references, and implicatures. Understanding such nuances is a requirement for successful natural…
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
Sourjyadip Ray, Shubham Sharma, Somak Aditya +1
As digital platforms redefine educational paradigms, ensuring interactivity remains vital for effective learning. This paper explores using Multimodal Large Language Models (MLLMs)…
TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
Sayantan Adak, Daivik Agrawal, Animesh Mukherjee +1
We investigate the knowledge of object affordances in pre-trained language models (LMs) and pre-trained Vision-Language models (VLMs). A growing body of literature shows that PTLMs…