logo Idiap Research Institute        
 [BibTeX] [Marc21]
Reference-based vs. task-based evaluation of human language technology
Type of publication: Conference paper
Citation: Popescu-Belis_LREC_2008
Booktitle: LREC 2008 ELRA Workshop on Evaluation
Year: 2008
Location: Marrakech, Morocco
Organization: ELRA
Abstract: This paper starts from the ISO distinction of three types of evaluation procedures – internal, external and in use – and proposes to match these types to the three types of human language technology (HLT) systems: analysis, generation, and interactive. The paper explains why internal evaluation is not suitable to measure the qualities of HLT systems, and shows that reference-based external evaluation is best adapted to ‘analysis’ systems, task-based evaluation to ‘interactive’ systems, while ‘generation’ systems can be subject to both types of evaluation. In particular, some limits of reference-based external evaluation are shown in the case of generation systems. Finally, the paper shows that contextual evaluation, as illustrated by the FEMTI framework for MT evaluation, is an effective method for getting reference-based evaluation closer to the users of a system.
Keywords:
Projects Idiap
IM2
Authors Popescu-Belis, Andrei
Added by: [UNK]
Total mark: 0
Attachments
  • Popescu-Belis_LREC_2008.pdf
Notes