Conference Proceedings

unimelb: Spanish Text Normalisation

B HAN, P Cook, TJ Baldwin

Proceedings of the Conference of the Spanish Society for Natural Language Processing | SEPLN (Sociedad Española para el Procesamiento del Lenguaje Natural) | Published : 2013

Abstract

This paper describes a lexicon-based text normalisation approach for Spanish tweets. We first compare English and Spanish text normalisation, and hypothesise that an approach previously proposed for English can be adapted to Spanish. A corpus-derived normalisation lexicon is built using distributional similarity, and is combined with existing lexicons (e.g., containing Spanish Internet slang). These lexicons enable a very fast, look-up based approach to text normalisation. Experimental results indicate that the corpus-derived lexicon complements existing lexicons, but that the approach could be improved through better handling of certain word types, such as named entities.

University of Melbourne Researchers

Citation metrics