2 papers
cs.CL2024
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Oggi Rudovic, Sameer Dharur +7
Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data. Incorporating audio or video improves performanc…
cs.LG2024
Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
Utkarsh Oggy Sarawgi, John Berkowitz, Vineet Garg +5
Streaming neural network models for fast frame-wise responses to various speech and sensory signals are widely adopted on resource-constrained platforms. Hence, increasing the lear…