2 papers
eess.AS2026
Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning
Mohan Li, Rama Doddipatla, Philip C. Woodland
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its a…
cs.CL2025
Conditional Multi-Stage Failure Recovery for Embodied Agents
Youmna Farag, Svetlana Stoyanchev, Mohan Li +2
Embodied agents performing complex tasks are susceptible to execution failures, motivating the need for effective failure recovery mechanisms. In this work, we introduce a conditio…