9 papers
MULTISEISMO: A Multimodal Seismic Dataset and Model for Cross-Modal Seismic Understanding
Sai Munikoti, Ian Stewart, Chengping Chai +4
The application of generalist multimodal models (GMMs) to specialized scientific domains remains limited due to the scarcity of comprehensive domain-specific datasets that integrat…
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
Sameera Horawalavithana, Lauren Phillips, Ian Stewart +2
Vision-Language Models (VLMs) have rapidly advanced by leveraging powerful pre-trained Large Language Models (LLMs) as core reasoning backbones. As new and more capable LLMs emerge…
SCITUNE: Aligning Large Language Models with Human-Curated Scientific Multimodal Instructions
Sameera Horawalavithana, Sai Munikoti, Ian Stewart +2
Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving LLMs to align e…
Directional Concentration Uncertainty: A representational approach to uncertainty quantification for generative models
Souradeep Chattopadhyay, Brendan Kennedy, Sai Munikoti +2
In the critical task of making generative models trustworthy and robust, methods for Uncertainty Quantification (UQ) have begun to show encouraging potential. However, many of thes…
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
Ian Stewart, Sameera Horawalavithana, Brendan Kennedy +2
Multimodal foundation models (MFMs) such as OFASys show the potential to unlock analysis of complex data such as images, videos, and audio data via text prompts alone. However, the…
Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing
Reilly Raab, Mike Parker, Dan Nally +4
The advent of language models (LMs) has the potential to dramatically accelerate tasks that may be cast to text-processing; however, real-world adoption is hindered by concerns reg…