Conference Proceedings

Crowd-Sourcing of Human Judgments of Machine Translation Fluency

Y Graham, TJ Baldwin, A Moffat, J Zobel

Proceedings of the Australasian Language Technology Association Workshop (ALTA) | ACL Anthology | Published : 2013

Abstract

Human evaluation of machine translation quality is a key element in the development of machine translation systems, as automatic metrics are validated through correlation with human judgment. However, achievement of consistent human judgments of machine translation is not easy, with decreasing levels of consistency reported in annual evaluation campaigns. In this paper we describe experiences gained during the collection of human judgments of the fluency of machine translation output using Amazon's Mechanical Turk service. We gathered a large collection of crowd-sourced human judgments for the machine translation systems that participated in the WMT 2012 shared translation task, collected ac..

View full abstract

Grants

Citation metrics