Conference Proceedings
How Well Do Embedding Models Capture Non-compositionality? A View from Multiword Expressions
Navnita Nandakumar, Timothy Baldwin, Bahar Salehi
Proceedings of the 3rd Workshop on Evaluating Vector Space Representations for | Association for Computational Linguistics | Published : 2019
DOI: 10.18653/v1/w19-2004
Open access
Abstract
In this paper, we apply various embedding methods on multiword expressions to study how well they capture the nuances of non-compositional data. Our results from a pool of word-, character-, and document-level embbedings suggest that Word2vec performs the best, followed by FastText and Infersent. Moreover, we find that recently-proposed contextualised embedding models such as Bert and ELMo are not adept at handling non-compositionality in multiword expressions.