Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications
Piyush Jain, Kousik Dasgupta, Rajarshi Roy +1
As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spat…
cs.CV2025
PALADIN : Robust Neural Fingerprinting for Text-to-Image Diffusion Models
Murthy L, Subarna Tripathi
The risk of misusing text-to-image generative models for malicious uses, especially due to the open-source development of such models, has become a serious concern. As a risk mitig…
cs.CV2025
ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way
Rajarshi Roy, Devleena Das, Ankesh Banerjee +3
We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Laye…