MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark

Haoran Li; Abhinav Arora; Shuohui Chen; Anchit Gupta; Sonal Gupta; Yashar Mehdad

doi:10.18653/v1/2021.eacl-main.257

MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark

Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, Yashar Mehdad

Abstract

Scaling semantic parsing models for task-oriented dialog systems to new languages is often expensive and time-consuming due to the lack of available datasets. Available datasets suffer from several shortcomings: a) they contain few languages b) they contain small amounts of labeled examples per language c) they are based on the simple intent and slot detection paradigm for non-compositional queries. In this paper, we present a new multilingual dataset, called MTOP, comprising of 100k annotated utterances in 6 languages across 11 domains. We use this dataset and other publicly available datasets to conduct a comprehensive benchmarking study on using various state-of-the-art multilingual pre-trained models for task-oriented semantic parsing. We achieve an average improvement of +6.3 points on Slot F1 for the two existing multilingual datasets, over best results reported in their experiments. Furthermore, we demonstrate strong zero-shot performance using pre-trained models combined with automatic translation and alignment, and a proposed distant supervision method to reduce the noise in slot label projection.

Anthology ID:: 2021.eacl-main.257
Volume:: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume
Month:: April
Year:: 2021
Address:: Online
Editors:: Paola Merlo, Jorg Tiedemann, Reut Tsarfaty
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 2950–2962
Language:
URL:: https://aclanthology.org/2021.eacl-main.257/
DOI:: 10.18653/v1/2021.eacl-main.257
Bibkey:
Cite (ACL):: Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021. MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 2950–2962, Online. Association for Computational Linguistics.
Cite (Informal):: MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark (Li et al., EACL 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.eacl-main.257.pdf
Data: MTOP

PDF Cite Search Fix data