Relational Programming with Foundation Models
arXiv:2412.14515 · doi:10.1609/aaai.v38i9.28934
Abstract
Foundation models have vast potential to enable diverse AI applications. The powerful yet incomplete nature of these models has spurred a wide range of mechanisms to augment them with capabilities such as in-context learning, information retrieval, and code interpreting. We propose Vieira, a declarative framework that unifies these mechanisms in a general solution for programming with foundation models. Vieira follows a probabilistic relational paradigm and treats foundation models as stateless functions with relational inputs and outputs. It supports neuro-symbolic applications by enabling the seamless combination of such models with logic programs, as well as complex, multi-modal applications by streamlining the composition of diverse sub-models. We implement Vieira by extending the Scallop compiler with a foreign interface that supports foundation models as plugins. We implement plugins for 12 foundation models including GPT, CLIP, and SAM. We evaluate Vieira on 9 challenging tasks that span language, vision, and structured and vector databases. Our evaluation shows that programs in Vieira are concise, can incorporate modern foundation models, and have comparable or better accuracy than competitive baselines.
References in corpus (25)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- On the Opportunities and Risks of Foundation Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Large Language Models are Zero-Shot Reasoners
- High-Resolution Image Synthesis with Latent Diffusion Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Segment Anything
- ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
- MPNet: Masked and Permuted Pre-training for Language Understanding
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding
- Prompting Is Programming: A Query Language for Large Language Models
- PAL: Program-aided Language Models
- Simple Open-Vocabulary Object Detection with Vision Transformers
- SATNet: Bridging deep learning and logical reasoning using a differentiable satisfiability solver
- Learning Reasoning Strategies in End-to-End Differentiable Proving
- Reliable Natural Language Understanding with Large Language Models and Answer Set Programming
- Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search
- Closed Loop Neural-Symbolic Learning via Integrating Neural Perception, Grammar Parsing, and Symbolic Reasoning
- Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems
- Plug-and-Play VQA: Zero-shot VQA by Conjoining Large Pretrained Models with Zero Training