Conference Proceedings
UniMelb_NLP-CORE: Integrating predictions from multiple domains and feature sets for estimating semantic textual similarity
S GELLA, B SALEHI, M LUI, K GRIESER, P Cook, TJ Baldwin
Proceedings of the 2nd Joint Conference on Lexical and Computational Semantics | Omnipress | Published : 2013
Abstract
In this paper we present our systems for calculating the degree of semantic similarity between two texts that we submitted to the Semantic Textual Similarity task at SemEval-2013. Our systems predict similarity using a regression over features based on the following sources of information: string similarity, topic distributions of the texts based on latent Dirichlet allocation, and similarity between the documents returned by an information retrieval engine when the target texts are used as queries. We also explore methods for integrating predictions using different training datasets and feature sets. Our best system was ranked 17th out of 89 participating systems. In our post-task analysis,..
View full abstract