4 papers
WhisperRT -- Turning Whisper into a Causal Streaming Model
Tomer Krichli, Bhiksha Raj, Joseph Keshet
Automatic Speech Recognition (ASR) has seen remarkable progress, with models like OpenAI Whisper and NVIDIA Canary achieving state-of-the-art (SOTA) performance in offline transcri…
Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
Thanh-Dat Truong, Huu-Thien Tran, Tran Thai Son +2
Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some f…
Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints
Yuxuan Wu, Ziyu Wang, Bhiksha Raj +1
We contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms t…
FLAASH: Flow-Attention Adaptive Semantic Hierarchical Fusion for Multi-Modal Tobacco Content Analysis
Naga VS Raviteja Chappa, Page Daniel Dobbs, Bhiksha Raj +1
The proliferation of tobacco-related content on social media platforms poses significant challenges for public health monitoring and intervention. This paper introduces a novel mul…