activity
20222024
most citedMoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

4 citations · 6 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CL2024

CoDi: Conversational Distillation for Grounded Question Answering

Patrick Huber, Arash Einolghozati, Rylan Conway +6

Distilling conversational skills into Small Language Models (SLMs) with approximately 1 billion parameters presents significant challenges. Firstly, SLMs have limited capacity in t…

cs.AI20244 cited

MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Xi Victoria Lin, Akshat Shrivastava, Liang Luo +5

We introduce MoMa, a novel modality-aware mixture-of-experts (MoE) architecture designed for pre-training mixed-modal, early-fusion language models. MoMa processes images and text…

cs.CL2024

PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding

Trang Le, Daniel Lazar, Suyoun Kim +6

Spoken Language Understanding (SLU) is a critical component of voice assistants; it consists of converting speech to semantic parses for task execution. Previous works have explore…

cs.CL20241 cited

Small But Funny: A Feedback-Driven Approach to Humor Distillation

Sahithya Ravi, Patrick Huber, Akshat Shrivastava +4

The emergence of Large Language Models (LLMs) has brought to light promising language generation capabilities, particularly in performing tasks like complex reasoning and creative…

cs.CL2023

Augmenting text for spoken language understanding with Large Language Models

Roshan Sharma, Suyoun Kim, Daniel Lazar +7

Spoken semantic parsing (SSP) involves generating machine-comprehensible parses from input speech. Training robust models for existing application domains represented in training d…

cs.CL2023

Modality Confidence Aware Training for Robust End-to-End Spoken Language Understanding

Suyoun Kim, Akshat Shrivastava, Duc Le +3

End-to-end (E2E) spoken language understanding (SLU) systems that generate a semantic parse from speech have become more promising recently. This approach uses a single model that…