2 papers
cs.CV2025
LLMs Can Compensate for Deficiencies in Visual Representations
Sho Takishita, Jay Gala, Abdelrahman Mohamed +2
Many vision-language models (VLMs) that prove very effective at a range of multimodal task, build on CLIP-based vision encoders, which are known to have various limitations. We inv…
cs.LG2025
SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations
Gaurav Sarkar, Jay Gala, Subarna Tripathi
The design of activation functions remains a pivotal component in optimizing deep neural networks. While prevailing choices like Swish and GELU demonstrate considerable efficacy, t…