Validation of an Automatic Metric for the Accuracy of Pronoun Translation (APT)
Type of publication: Idiap-RR
Citation: Werlen_Idiap-RR-29-2016
Number: Idiap-RR-29-2016
Year: 2016
Month: 11
Institution: Idiap
Abstract: In this paper, we define and assess a reference-based metric to evaluate the accuracy of pronoun translation (APT). The metric automatically aligns a candidate and a reference translation using GIZA++ augmented with specific heuristics, and then counts the number of identical or different pronouns, with provision for legitimate variations and omitted pronouns. All counts are then combined into one score. The metric is applied to the results of seven systems (including the baseline) that participated in the DiscoMT 2015 shared task on pronoun translation from English to French. The APT metric reaches around 0.993-0.999 Pearson correlation with human judges (depending on the parameters of APT), while other automatic metrics such as BLEU, METEOR, or those specific to pronouns used at DiscoMT 2015 reach only 0.972-0.986 Pearson correlation.
Projects Idiap
Authors Miculicich, Lesly
Popescu-Belis, Andrei
Crossref by MiculicichWerlen_DISCOMTATEMNLP_2017
  Werlen_Idiap-RR-29-2016.pdf