Measuring machine translation quality as semantic equivalence: A metric based on entailment features

详细信息

作者：Sebastian Pad贸 ; Daniel Cer ; Michel Galley ; Dan Jurafsky and Christopher D. Manning
关键词：MT evaluation ; Automated metric ; MERT ; Semantics ; Entailment ; Linguistic analysis ; Paraphrase
刊名：Machine Translation
出版时间：September, 2009
出版年：2009
期刊代码：36_09226567
类别：cp
卷：23
期：2-3
页码：181-193
数据来源：sp

摘要

Current evaluation metrics for machine translation have increasing difficulty in distinguishing good from merely fair translations. We believe the main problem to be their inability to properly capture meaning: A good translation candidate means the same thing as the reference translation, regardless of formulation. We propose a metric that assesses the quality of MT output through its semantic equivalence to the reference translation, based on a rich set of match and mismatch features motivated by textual entailment. We first evaluate this metric in an evaluation setting against a combination metric of four state-of-the-art scores. Our metric predicts human judgments better than the combination metric. Combining the entailment and traditional features yields further improvements. Then, we demonstrate that the entailment metric can also be used as learning criterion in minimum error rate training (MERT) to improve parameter estimation in MT system training. A manual evaluation of the resulting translations indicates that the new model obtains a significant improvement in translation quality.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700