Abstract
We have recently seen the emergence of several publicly available Natural Language Understanding (NLU) toolkits, which map user utterances to structured, but more abstract, Dialogue Act (DA) or Intent specifications, while making this process accessible to the lay developer. In this paper, we present the first wide coverage evaluation and comparison of some of the most popular NLU services, on a large, multi-domain (21 domains) dataset of 25 K user utterances that we have collected and annotated with Intent and Entity Type specifications and which will be released as part of this submission (https://github.com/xliuhw/NLU-Evaluation-Data ). The results show that on Intent classification Watson significantly outperforms the other platforms, namely, Dialogflow, LUIS and Rasa; though these also perform well. Interestingly, on Entity Type recognition, Watson performs significantly worse due to its low Precision (At the time of producing the camera-ready version of this paper, we noticed the seemingly recent addition of a ‘Contextual Entity’ annotation tool to Watson, much like e.g. in Rasa. We’d threfore like to stress that this paper does not include an evaluation of this feature in Watson NLU.). Again, Dialogflow, LUIS and Rasa perform well on this task.
Original language | English |
---|---|
Title of host publication | Increasing Naturalness and Flexibility in Spoken Dialogue Interaction |
Subtitle of host publication | 10th International Workshop on Spoken Dialogue Systems |
Publisher | Springer |
Pages | 165-183 |
Number of pages | 19 |
Edition | 1 |
ISBN (Electronic) | 9789811593239 |
ISBN (Print) | 9789811593222, 9789811593253 |
DOIs | |
Publication status | Published - 2021 |
Event | 10th International Workshop on Spoken Dialogue Systems Technology 2019 - Sicily, Siracusa, Italy Duration: 24 Apr 2019 → 26 Apr 2019 https://iwsds2019.unikore.it/ |
Publication series
Name | Lecture Notes in Electrical Engineering |
---|---|
Volume | 714 |
ISSN (Print) | 1876-1100 |
ISSN (Electronic) | 1876-1119 |
Conference
Conference | 10th International Workshop on Spoken Dialogue Systems Technology 2019 |
---|---|
Abbreviated title | IWSDS 2019 |
Country/Territory | Italy |
City | Siracusa |
Period | 24/04/19 → 26/04/19 |
Internet address |
ASJC Scopus subject areas
- Industrial and Manufacturing Engineering
Fingerprint
Dive into the research topics of 'Benchmarking Natural Language Understanding Services for Building Conversational Agents'. Together they form a unique fingerprint.Datasets
-
NLU benchmark
Rieser, V. (Creator) & Liu, X. (Creator), Heriot-Watt University, Feb 2019
https://github.com/xliuhw/NLU-Evaluation-Data/
Dataset