3 papers
cs.LG2026
Tactus: Open-Vocabulary Object Recognition from Low-Cost Pressure Arrays
Abdul Basit Tonmoy
Resistive pressure arrays are the cheapest and most widely shipped tactile sensors, yet tactile representation learning has concentrated on optical sensors that image a deforming g…
cs.CL2026
Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding
Abdul Basit Tonmoy
Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-spee…
cs.CL2026
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Abdul Basit Tonmoy, Kazi Fardinul Hoque, Md. Shahrier Islam Arham +1
A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead t…