From the 1 of 4 linked papers with an AI index.
4 papers
Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications
Piyush Jain, Kousik Dasgupta, Rajarshi Roy +1
The paper introduces ByDeWay-V2, a training‑free prompting framework that adds explicit pairwise spatial predicates derived from depth estimation and open‑vocabulary object detecti…
ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way
Rajarshi Roy, Devleena Das, Ankesh Banerjee +3
We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Laye…
PALADIN : Robust Neural Fingerprinting for Text-to-Image Diffusion Models
Murthy L, Subarna Tripathi
The risk of misusing text-to-image generative models for malicious uses, especially due to the open-source development of such models, has become a serious concern. As a risk mitig…
SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
Gaurav Sarkar, Jay Gala, Subarna Tripathi
The design of activation functions remains a pivotal component in optimizing deep neural networks. While prevailing choices like Swish and GELU demonstrate considerable efficacy, t…