mRAT-SQL+GAP:A Portuguese Text-to-SQL Transformer
arXiv:2110.03546 · doi:10.1007/978-3-030-91699-2_35
Abstract
The translation of natural language questions to SQL queries has attracted growing attention, in particular in connection with transformers and similar language models. A large number of techniques are geared towards the English language; in this work, we thus investigated translation to SQL when input questions are given in the Portuguese language. To do so, we properly adapted state-of-the-art tools and resources. We changed the RAT-SQL+GAP system by relying on a multilingual BART model (we report tests with other language models), and we produced a translated version of the Spider dataset. Our experiments expose interesting phenomena that arise when non-English languages are targeted; in particular, it is better to train with original and translated training datasets together, even if a single target language is desired. This multilingual BART model fine-tuned with a double-size training dataset (English and Portuguese) achieved 83% of the baseline, making inferences for the Portuguese test dataset. This investigation can help other researchers to produce results in Machine Learning in a language different from English. Our multilingual ready version of RAT-SQL+GAP and the data are available, open-sourced as mRAT-SQL+GAP at: https://github.com/C4AI/gap-text2sql
Published in: Intelligent Systems. BRACIS 2021. Lecture Notes in Computer Science
References in corpus (13)
- Learning to Map Sentences to Logical Form: Structured Classification with Probabilistic Categorial Grammars
- Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
- SQLNet: Generating Structured Queries From Natural Language Without Reinforcement Learning
- A Comparative Survey of Recent Natural Language Interfaces for Databases
- Multilingual Translation with Extensible Multilingual Pretraining and Finetuning
- GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
- Hybrid Ranking Network for Text-to-SQL
- Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate Representation
- Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
- Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing
- Duoquest: A Dual-Specification System for Expressive SQL Queries
- Semantic Evaluation for Text-to-SQL with Distilled Test Suites
- LGESQL: Line Graph Enhanced Text-to-SQL Model with Mixed Local and Non-Local Relations