Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding
Abdul Basit Tonmoy
Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-spee…
cs.CL2026
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Abdul Basit Tonmoy, Kazi Fardinul Hoque, Md. Shahrier Islam Arham +1
A single embedding space that covers text, images, video, and audio lets one index serve every query a user can pose. Embedding models built on vision-language backbones now lead t…